Which company has the best AI model end of September?
Anthropic is still the favorite to hold the top Arena rank at the end of September, but the market may be a bit too confident given how quickly leaderboard positions can change. I would price Yes lower than the current market and put it at 84%.
Analysis
Anthropic appears well positioned because this market is about the company whose model sits at the top of the Arena text leaderboard at a single check time, not about long-run model quality in the abstract. In that setting, a company that is already near the front of the pack usually has a meaningful advantage, especially when the remaining time horizon is only about one month. Anthropic’s models have historically been among the strongest in user-preference style evaluations, which tends to fit this leaderboard format well.
At the same time, the Arena ranking is notoriously dynamic. A new release, a quiet model refresh, a change in sampling behavior, or even shifts in how the community interacts with the interface can move the order quickly, and the margin between first and second is often small enough that granular score differences matter. Because the market resolves to the company in first place at a single timestamp, Anthropic does not need to be broadly superior in a general sense, only narrowly ahead on that day, which makes this much more fragile than a typical long-duration fundamental bet.
The market price of 92% suggests participants believe Anthropic is either already ahead by a comfortable margin or is highly likely to remain there through the end of the month. I think that confidence is somewhat inflated. Anthropic is the most plausible winner, but the combination of short horizon, leaderboard volatility, and the possibility of a competitor shipping a strong update means the true probability should be high but not overwhelming. A fair estimate is that Anthropic remains the best model at the September check more often than not, though not as close to certain as the market implies.
Arguments
For
- Arguments for Yes: Anthropic is a strong fit for user-preference text leaderboards and has frequently been among the top contenders.
- Arguments for Yes: The short remaining window makes it harder for a rival to mount a sustained overtaking run before the check date.
Against
- Arguments against Yes: Frontier model leadership is unstable, and a single strong release from a rival can reverse the order quickly.
- Arguments against Yes: The leaderboard is checked at one moment and exact score tie-breaks can decide the outcome even when models are effectively neck and neck.
Key drivers
- Anthropic has a strong track record in preference-based text evaluations, which aligns well with Arena ranking dynamics.
- Only about one month remains, so the current leader has less time to be overtaken by a competitor.
- The market resolves on a single timestamp, which favors the incumbent if the lead is already established.
- A new release from a rival company could quickly change the leaderboard before the September 30 check.
Risk factors
- Leaderboard ranks can shift quickly because small score changes or tie-breaks can flip first place.
- OpenAI, Google, or xAI could launch or refresh a model that overtakes Anthropic before month-end.
- Arena preference is sensitive to sampling, prompt mix, and community behavior, which adds noise to the ranking.
- The market is already pricing a very high probability, leaving limited room for upside if conditions do not improve further.
Scenarios
Best case
Anthropic keeps a clear lead or benefits from a small score edge at the September 30 check, making it the top-ranked company at resolution.
Most likely
Anthropic remains one of the top two or three contenders, but there is enough leaderboard volatility that a last-minute shift from another frontier model is still plausible.
Worst case
A competitor such as OpenAI or Google ships a stronger model or update that overtakes Anthropic before the check, pushing Anthropic out of first place.
More from this day
- techPolymarketEnded
Next Claude Opus released by...?
AI97%MKT9%Edge+88Hidden GemThe balance of evidence strongly favors Yes because the latest Claude Opus release already appears to have occurred before the deadline. The main uncertainty is not whether a release happened, but whether the market ultimately interprets the wording in an unusually narrow way.
- pop culturePolymarketEnded
"Spider-Man: Brand New Day" 5th Weekend Box Office
AI78%MKT5%Edge+73Hidden GemThe current evidence leans strongly toward a fifth-weekend gross below $20 million, with the most relevant recent projections clustering around $18 million. The main reason not to go even higher is that these are still estimates rather than final figures, and a modest upside surprise could keep it above the threshold.
- CompaniesKalshi1y
Robinhood funded customers in 2026
AI44%MKT19%Edge+25Hidden GemRobinhood has a plausible path to exceed 30.2 million funded customers in 2026, but the hurdle is high enough that the outcome is still closer to a coin flip than the market implies. My independent estimate is meaningfully above the market price, though not enough to call it a strong Yes.