Which company has the best AI model on LiveBench (Coding) end of September?
Anthropic looks like the leading candidate to finish first on LiveBench Coding at the end of September, because it is already on top in the latest snapshot and the market is strongly aligned with that view. Still, the outcome is not locked in, since a late-month model release from a rival could change the leaderboard before the final check.
Analysis
Anthropic is currently the most likely winner because it already appears to hold the top position on the LiveBench Coding leaderboard in the latest available snapshot. For a market like this, the incumbent leader has a meaningful advantage: if the benchmark is only checked once at month-end, the company that is already first needs to simply avoid being displaced, while competitors must actually overtake it. That makes the current setup favorable to Yes, especially when the public market is already pricing a strong probability in that direction.
The strongest argument for Anthropic is that its recent model family appears to be very competitive specifically on coding tasks, which is exactly what determines this market. If the current leader remains stable through the end of September, the company benefits from both benchmark momentum and the practical difficulty of large jumps in ranking over a short time frame. In addition, benchmark leaderboards often move in small increments rather than dramatic swings, so an established first-place model can remain ahead unless another lab ships something clearly better very late in the month.
The main reason not to assign an even higher probability is that the market is still exposed to late-release risk. A rival company could publish a stronger coding model before September 30, and because the market resolves strictly by the end-of-month leaderboard snapshot, one new release is enough to overturn the current ranking. There is also some uncertainty in how stable the leaderboard will be, whether the score differences are narrow, and whether any tie-breakers become relevant if multiple models cluster tightly near the top. That means the current leader has a clear edge, but not an overwhelming one.
Market sentiment reinforces the Yes side, since the contract is trading well above 80%, which suggests traders broadly believe Anthropic will keep the lead. However, market prices can overstate certainty when an incumbent is visible on the board and the remaining time window is short. My independent assessment is slightly below the market price because the resolution depends on a specific future snapshot and because frontier-model competition is still capable of producing a late upset.
Arguments
For
- Arguments for Yes: Anthropic is already listed as the top Coding model in the latest leaderboard snapshot, so it starts from first place.
- Arguments for Yes: The market is heavily tilted toward Anthropic, which suggests informed traders see the current lead as durable.
Against
- Arguments against Yes: Another lab can still launch a better coding model before the end-of-month check and take first place.
- Arguments against Yes: The resolution is based on one specific leaderboard snapshot, so even a brief late change would be enough to invalidate Anthropic's lead.
Key drivers
- Anthropic is already the current LiveBench Coding leader, which gives it a direct advantage heading into the final snapshot.
- The market is strongly priced toward Yes, indicating broad expectation that the lead will hold through month-end.
- Late September model launches from competitors could still overtake the current score and flip the ranking.
- The final outcome depends on a single end-of-month check, so small leaderboard changes matter a lot.
Risk factors
- A competitor may release a materially stronger coding model before September 30 and surpass Anthropic.
- If the top scores are very close, a small benchmark update or tie-break rule could change the winner unexpectedly.
Scenarios
Best case
Anthropic keeps the top Coding score through September 30, no competitor ships a clearly better model, and any ties are resolved in its favor or never become relevant.
Most likely
Anthropic remains near the top and probably finishes first, but the market still has to survive the risk of a last-minute competitor release before the final leaderboard snapshot.
Worst case
A rival model improves enough in late September to pass Anthropic on LiveBench Coding, pushing Anthropic out of first place at the final check.
More from this day
- PoliticsKalshi7d
When will a reconciliation bill become law?
AI99%MKT2%Edge+97Hidden GemA reconciliation bill appears to have already become law, which would satisfy the condition well before Oct. 1, 2026. On the facts provided, I put the Yes probability extremely high unless the market uses an unusually narrow resolution rule tied to a different bill.
- PoliticsKalshi7d
When will Trump nominate a Federal Reserve governor?
AI99%MKT3%Edge+96Hidden GemA nomination has already been made, so the Yes outcome is overwhelmingly likely unless the market is using an unusually narrow definition of nomination. The current price looks like a clear misread of the event timing and the underlying news flow.
- techPolymarket9mo
Next Grok Model (4.7+): Text Arena Debut?
AI88%MKT39%Edge+49Hidden GemThe evidence strongly suggests the next qualifying Grok model has already debuted and that its public launch scores were well above the 1450 threshold, so the Yes case is very strong. The main uncertainty is whether the market’s exact leaderboard-and-timing rules are satisfied by the available reporting, not whether the model itself is capable of reaching the score.