Which company has the best AI model end of September?
Anthropic is the current favorite because Claude models are already near the top of the relevant arena leaderboard and recent benchmark commentary is consistently strong. Still, the contract is more fragile than the market price suggests because a single competitor release or leaderboard swing could flip the ranking before the end of September.
Analysis
Anthropic enters this market with a clear near-term advantage. Recent landscape summaries repeatedly place Claude Opus 5 and Claude Fable 5 at or near the top across independent evaluations, and at least one current arena-style leaderboard snapshot has Claude Fable 5 in first place. Since the resolution depends on the specific arena.ai Text Arena Overall ranking on September 30, the most relevant fact is not just broad model quality but whether Anthropic is already winning the exact leaderboard that will be checked, and the available context suggests it is currently in that position or very close to it.
The main reason to discount a near-certainty outcome is that leaderboard leadership in this space can change quickly. OpenAI’s newest frontier models are described as close behind, xAI’s top model is also in the mix, and the gap between first and second may be small enough that a single new release, an aggressive tuning update, or a preference-shift in arena voting could move the top slot. Because the market resolves on a snapshot rather than an average over time, Anthropic does not need to be the best model in some general sense; it only needs to be first at one specific check point, which raises volatility rather than reducing it.
The market price of 87.5% implies strong confidence that Anthropic will still be on top, but I think that slightly overstates the true probability. Anthropic’s position is genuinely strong, and the company has multiple frontier models that could defend the lead, but the remaining six weeks leave enough room for a rival jump or an Anthropic slip to make this a meaningful upset risk. My assessment is that Yes is more likely than not by a wide margin, yet not so dominant that the price should be treated as nearly locked in.
Arguments
For
- Arguments for Yes: Anthropic is already reported to be leading or near-leading on the exact type of arena ranking that matters for resolution.
- Arguments for Yes: Claude Opus 5 and Claude Fable 5 give Anthropic two strong chances to hold the top slot even if one model is overtaken.
Against
- Arguments against Yes: The leaderboard is competitive and recent summaries say OpenAI and xAI are close enough to challenge the lead quickly.
- Arguments against Yes: A new release or a small leaderboard score shift could easily change first place by the end-of-month check.
Key drivers
- Anthropic currently appears to have the strongest or tied-strongest models on the relevant arena leaderboard.
- The resolution depends on a single ranking snapshot, which amplifies short-term leaderboard volatility.
Risk factors
- OpenAI or xAI could launch or promote a model that overtakes Anthropic before September 30.
- Small score differences or tie-break rules could flip first place even if model quality remains broadly similar.
Scenarios
Best case
Anthropic keeps the top arena rank through September because its current lead holds and no competitor ships a clearly superior model before the check date.
Most likely
Anthropic remains one of the top two contenders and has a better than even chance to finish first, but the final outcome is still sensitive to late model launches and small ranking changes.
Worst case
A rival model from OpenAI or another major lab overtakes Anthropic on the arena leaderboard in late September, pushing Anthropic out of first place at the resolution time.
More from this day
- EconomicsKalshi9y
US real GDP growth in 2035?
AI79%MKT15%Edge+64Hidden GemMy independent view is that U.S. real GDP growth in 2035 is more likely to land in the 1.6% to 2.5% range than to be at or below zero, with a meaningful but not dominant chance of stronger growth. The market looks too pessimistic on the downside and slightly underweights a broad middle case of moderate expansion.
- techPolymarketEnded
Best Chinese AI Company end of August?
AI43%MKT96%Edge-53HypedAlibaba is not currently first on the relevant Chinese-model leaderboard, so the market looks too optimistic. A late-month Qwen update could still flip the standings, but that is more plausible than likely.
- pop culturePolymarketEnded
"Coyote vs Acme" Rotten Tomatoes Score?
AI47%MKT96%Edge-49HypedI think the market is far too confident. Coyote vs Acme has a plausible path to a strong Rotten Tomatoes score, but the 85 threshold is high enough that the more realistic outcome is a finish below it or, less likely, no score available by the deadline.