Next Mythos-Class Model: Text Arena Debut?
Anthropic’s next Mythos-class model looks more likely than not to clear 1470 if it appears publicly on the Arena leaderboard, but the market is pricing in too much confidence because visibility and leaderboard qualification are still meaningful hurdles. I would put the chance of a Yes at 71%.
Analysis
The main reason to lean Yes is that a next-generation Anthropic frontier model should be competitive enough to land in the high 1400s if it receives a normal public debut on the text leaderboard. The recent context suggests the Mythos line sits above Opus, and the broader benchmark picture shows frontier models already clustering near the threshold region, so 1470 is not an extreme bar for a model in this tier. The threshold is high, but it is still inside the range where a top-tier release can plausibly clear it on first appearance.
The biggest reason to be more cautious than the market is the resolution setup itself. This market does not simply ask whether the model is strong; it requires a qualifying, non-AutoEval, publicly visible leaderboard debut, and the latest context suggests Anthropic’s newest Mythos-class model may not be broadly exposed in a way that reliably generates a valid Arena score. If the company keeps the model gated, releases it in limited form, or the leaderboard entry is delayed, removed, or labeled in a way that does not count, the market resolves No even if the model is excellent in private testing.
The current market price implies a very high confidence level, but that seems aggressive given the combined uncertainty around release timing, accessibility, and score variance. A frontier model can easily end up a bit below a round threshold like 1470 if the evaluation snapshot is noisy, if the system prompt or style-control conditions are unfavorable, or if the debut version is tuned conservatively rather than optimized for Arena. My estimate therefore stays above 50% because the model class is strong and the threshold is reachable, but meaningfully below the market because the public-debut condition is doing a lot of work.
Arguments
For
- Arguments for Yes: Anthropic’s frontier models are strong enough that a Mythos-class debut above 1470 is a realistic outcome.
- Arguments for Yes: The threshold is high but not exceptional relative to the best current text-arena results.
Against
- Arguments against Yes: The model may never appear in a qualifying public form, which would force a No resolution.
- Arguments against Yes: Even a strong debut can miss a tight cutoff like 1470 because Arena scores are noisy and sensitive to deployment details.
Key drivers
- Anthropic’s Mythos tier appears strong enough that a public debut above 1470 is plausible.
- The market depends on a qualifying public leaderboard entry, not just raw model quality.
- Score noise and debut configuration can easily move a frontier model a few points around the cutoff.
- A near-term release window would give the model time to appear before year-end 2026.
Risk factors
- Anthropic may not make the next Mythos-class model publicly visible in a way that counts for the leaderboard.
- The debut score could land just below 1470 even if the model is broadly competitive.
- The model could be labeled or handled in a way that makes it ineligible as an added leaderboard entry.
- Leaderboard volatility or removal after first appearance could prevent a valid resolution.
Scenarios
Best case
Anthropic releases the next Mythos-class model publicly, it appears on the leaderboard promptly, and it debuts comfortably above 1470, making the market resolve Yes without controversy.
Most likely
The model eventually appears publicly and is competitive, but the final outcome hinges on whether Anthropic chooses a qualifying release format and whether the first visible score clears the cutoff by a small margin.
Worst case
The model is kept too restricted for a valid leaderboard debut, or it appears but lands below 1470, causing the market to resolve No.
More from this day
- politicsPolymarket3mo
Google Maps renames Lake Ontario to "Lake America" by...?
AI5%MKT95%Edge-90HypedThis looks overwhelmingly unlikely. Google Maps would need to roll out a broad, production-level U.S. renaming of a major international lake within months, and there is no sign of that happening.
- pop culturePolymarketEnded
"Spider-Man: Brand New Day" 5th Weekend Box Office
AI90%MKT3%Edge+87Hidden GemThe under-20m outcome looks strongly favored for a Spider-Man film in its fifth weekend. Unless it is holding far better than typical superhero releases, the domestic weekend should be comfortably below the threshold.
- pop culturePolymarketEnded
"The Dog Stars" Opening Weekend Box Office
AI41%MKT93%Edge-52HypedThe under on 8 million is plausible, but the broader tracking picture still leans a bit above that line. I think the market is slightly overpricing the chance of a sub-8 million opening, so I put Yes at 41%.