Which company has the best AI model end of September?
Anthropic looks like the right favorite because the exact leaderboard used for settlement currently appears to favor Claude models and the competition is close but not clearly ahead. Still, the market price is a bit aggressive, so I would put the chance of Yes below the implied 90.5%.
Analysis
The most important point is that this market does not ask for the best model in some broad abstract sense; it asks for the company that owns the model sitting in first place on a very specific Text Arena leaderboard at a specific check time. On that exact kind of leaderboard, recent evidence suggests Anthropic is already extremely well positioned, with Claude models repeatedly appearing at or near the top and in some snapshots occupying the leading slot. That makes Yes the default favorite because the settlement rule is narrow and currently seems to align with Anthropic’s strongest area of performance.
The main reason not to make this a near-certain outcome is that frontier leaderboard leadership is fragile. The gap between the top models is described as very small, and recent comparisons show OpenAI, xAI, and Google all close enough that a single release, a scoring adjustment, or a modest quality jump can reshuffle the order. Since the market resolves at the end of September rather than immediately, there is still enough time for one of those competitors to launch or update a model that briefly or permanently takes first place on the exact arena ranking.
Market pricing leans heavily toward Anthropic, which is consistent with the public perception that Claude is currently the strongest all-around chat model on this specific benchmark family. However, the implied price is higher than I would personally assign because the market may be extrapolating today’s leaderboard into late September without fully discounting volatility. The right mental model is that Anthropic is favored, but not locked in; this is more like a high-probability lead than a done deal, and the remaining uncertainty mostly comes from the chance of a late competitive jump rather than from Anthropic’s current position being weak.
Arguments
For
- Arguments for Yes: Anthropic seems to have the strongest current position on the exact style-off Text Arena leaderboard that matters for resolution.
- Arguments for Yes: Public leaderboard momentum and prediction-market sentiment both currently point toward Anthropic as the most likely leader.
- Arguments for Yes: The benchmark setup favors models that perform well in broad conversational quality, which has been a strength of Claude models recently.
Against
- Arguments against Yes: The top models are clustered closely enough that a single update could flip first place before the check date.
- Arguments against Yes: Competitors such as OpenAI, xAI, and Google have enough research and release cadence to threaten a late change in rank.
- Arguments against Yes: The market may be overconfident because it is pricing recent Anthropic dominance as if it will persist unchanged through month-end.
Key drivers
- Anthropic currently appears to be leading or near-leading on the exact leaderboard format used for settlement.
- The current gap to rivals is narrow, so Anthropic only needs to avoid being overtaken rather than create a large lead.
- The remaining time until resolution is short enough to favor incumbency, but long enough for one major model update to matter.
Risk factors
- A late September release from OpenAI, xAI, or Google could quickly displace Anthropic on the specific arena ranking.
- Leaderboard rankings can move on small score changes, so first place is more fragile than the market price suggests.
Scenarios
Best case
Anthropic keeps or extends its current lead on the exact arena leaderboard, and one of the Claude models remains first at the September 30 check.
Most likely
Anthropic remains one of the top two or three companies on the leaderboard and likely stays in first place, but the margin stays tight enough that a late shuffle remains the main threat.
Worst case
A competitor launches or updates a model that edges Anthropic out of first place on the settlement leaderboard, making the answer No.
More from this day
- PoliticsKalshi2y
Which agencies will Trump eliminate?
AI82%MKT29%Edge+53Hidden GemI assess a high likelihood that USAID will be considered eliminated within the market’s timeframe, though the exact legal and institutional end state may continue to be contested. The strongest evidence points to a completed dismantling of the agency’s original structure and functions, with foreign aid largely absorbed into the State Department.
- cryptoPolymarket1y
Variational FDV above ___ one day after launch?
AI12%MKT58%Edge-46HypedThe market appears to be pricing this far too high relative to comparable launch-linked signals. My read is that Variational can launch successfully, but crossing an $800M FDV one day later looks unlikely unless the token debuts with very tight supply and unusually strong demand.
- EconomicsKalshi7y
US real GDP growth in 2033?
AI58%MKT12%Edge+46Hidden GemThe most likely outcome is that US real GDP growth in 2033 lands in the mid-range, with a meaningful but not dominant chance of a recessionary or very weak-growth year. My independent view is somewhat less bearish than the market on the top-range scenarios, and somewhat more supportive of moderate growth outcomes.