Which company has the best AI model on LiveBench (Coding) end of August?
Anthropic is the favorite, but the market looks a bit too confident for a leaderboard that can change quickly. I think Anthropic is more likely than not to finish first, though not at the extreme odds implied by the current price.
Analysis
The market is pricing Anthropic as a very strong favorite at 93.5%, which suggests either a current leadership position on the coding leaderboard or a perception that Claude’s coding strengths are hard to dislodge. That baseline is credible because Anthropic has been consistently competitive on programming and code-reasoning tasks, and LiveBench rewards broad model quality rather than a narrow trick that can be easily gamed. Still, a 93.5% implied chance leaves very little room for normal leaderboard volatility, which is unusually optimistic for an outcome determined by a live ranking snapshot near month-end.
Historically, coding leaderboards are among the most fluid benchmark categories. OpenAI and Google have both shown the ability to leapfrog competitors with new releases, tuning updates, or incremental model improvements, and Anthropic itself is not immune to being overtaken if another lab lands a stronger late-August model. Because the market resolves by whoever is first on the leaderboard at a specific check time, even a small temporary advantage or a recent model drop can matter more than long-run average performance. The exact tie-break rules also reduce the safety margin for a favorite if the top scores are close.
My independent estimate is lower than the market because there is still meaningful launch and update risk over the next few weeks. If Anthropic already has a clear lead and no major competitor release is imminent, Yes is the most likely outcome; however, if OpenAI or Google ships a stronger coding model before August 31, the top spot can change quickly. The best reading is that Anthropic remains the front-runner, but the true probability is closer to the low-80s than the mid-90s.
Arguments
For
- Arguments for Yes: Anthropic has one of the strongest reputations for code generation and coding reasoning, which makes it a natural favorite for this category.
- Arguments for Yes: If the leaderboard is already led by Anthropic or very close, it may be hard for rivals to catch up without a meaningful new model release.
Against
- Arguments against Yes: LiveBench is a live, competitive leaderboard, so a late release from a major rival could overturn the current order quickly.
- Arguments against Yes: The current price implies near-certainty, but benchmark rankings are often more fragile than that because small score changes can decide first place.
Key drivers
- Anthropic has a strong track record on coding and reasoning benchmarks, which supports a high baseline chance of leading LiveBench Coding.
- The market is already heavily tilted toward Yes, suggesting Anthropic may currently be near the top or perceived as the most stable leader.
- Leaderboard snapshots can be influenced by a single new release, so the company with the strongest present model often remains in front if no rival ships a major update.
- Tie-break rules based on cost per successful task can matter if scores are close, and Anthropic may benefit if its top score is both high and efficient.
Risk factors
- OpenAI or Google could release a better coding model before the end of August and immediately displace Anthropic.
- Small score differences on a live leaderboard can flip first place without any meaningful change in underlying model quality.
- If Anthropic’s best model is more expensive, the cost tie-breaker could hurt it in a close race.
- Benchmark-specific tuning by a rival could temporarily boost coding performance enough to win the month-end snapshot.
Scenarios
Best case
Anthropic keeps or widens its lead on the coding leaderboard, and no competitor ships a stronger model before the August 31 snapshot, so it clearly finishes first.
Most likely
Anthropic remains one of the top coding models and has a better-than-even chance to stay in first place, but the final outcome still depends on whether a rival update arrives before month-end.
Worst case
OpenAI or Google releases a superior coding model in late August and overtakes Anthropic, or Anthropic loses a close tie on score or cost efficiency.
More from this day
- FinancialsKalshi13y
Will OpenAI or Anthropic IPO first?
AI28%MKT92%Edge-64HypedAnthropic currently looks more likely to reach the public markets first, so I put OpenAI first at a fairly low probability. The market’s 89% implied yes price looks too optimistic unless OpenAI’s timetable has recently accelerated behind the scenes.
- pop culturePolymarketEnded
Kai and Speed finish their Minecraft marathon by...?
AI6%MKT62%Edge-56HypedMy read is that a Yes is possible but still unlikely, because the stream has to end cleanly before the deadline and remain ended for 12 hours. With no fresh confirmation that the marathon is wrapping up, the market’s heavy No price looks broadly justified.
- PoliticsKalshi2y
Which agencies will Trump eliminate?
AI84%MKT30%Edge+54Hidden GemUSAID looks much more likely than not to count as eliminated during Trump’s term, because the administration has already dismantled its independent operations and shifted its functions into the State Department. The main uncertainty is whether the market requires formal legal abolition, but even under that standard the path toward Yes remains materially stronger than the current price implies.