Highest Claude score on Humanity’s Last Exam in 2026?
Claude has a credible path to 55% on Humanity’s Last Exam by year-end, especially if Anthropic ships another strong frontier model or improves test-time reasoning. I think the chance is real but a bit below the current market price.
Analysis
Humanity’s Last Exam is an unusually demanding benchmark, so 55% is not a trivial milestone even for a frontier model family. The key detail is that this market is asking whether any Anthropic Claude model published in 2026 will clear the threshold, not whether the current model already has done so. That makes the question more favorable to Yes than a snapshot forecast, because Anthropic still has several months left to post a stronger public result and can benefit from both model improvements and evaluation refinements.
The strongest argument for Yes is the pace at which frontier models have been improving on hard reasoning and knowledge tasks. Anthropic has every incentive to keep Claude competitive, and a late-year release could plausibly add enough capability, better post-training, or stronger test-time computation to move a benchmark score by several points. If Claude is already in the general neighborhood of the low-to-mid 50s, then a modest generation jump, especially on a public leaderboard, is enough to push the family over the line. Because the market counts the highest score across all 2026 Claude entries, there is also more than one opportunity for Anthropic to succeed before the deadline.
The main reason to stay cautious is that HLE is designed to resist easy gains, so progress near the top may show diminishing returns. Moving from respectable to excellent may be harder than moving from weak to decent, and 55% may require a meaningful architectural or training advance rather than routine scaling. Anthropic may also prioritize safety, general product quality, or broader reasoning improvements over leaderboard chasing, which means a model that feels much better in practice might still miss this specific bar. On balance, the market’s 63.5% Yes price looks somewhat optimistic, but not wildly so; I would price the event just under that level because the threshold is reachable while still depending on a fairly strong late-2026 Claude release.
Arguments
For
- Arguments for Yes: Anthropic is likely to release at least one stronger Claude model before year-end, and a frontier jump could be enough to cross 55%.
- Arguments for Yes: Public benchmark optimization, including better reasoning and test-time compute, can sometimes add several points on difficult evaluations.
- Arguments for Yes: The market already expects a majority chance, which implies many traders think Claude is reasonably close to the target.
- Arguments for Yes: The market resolves on the highest 2026 score, so Anthropic has more than one opportunity to clear the bar.
Against
- Arguments against Yes: HLE is so difficult that 55% may require a substantial advance rather than normal iteration.
- Arguments against Yes: Anthropic could focus on product quality and safety over benchmark-specific tuning, limiting the public score.
- Arguments against Yes: If the best 2026 Claude result arrives late or not at all, there may be no time to recover from a miss.
- Arguments against Yes: Scores near the frontier often improve slowly, so the remaining gap could be larger than traders assume.
Key drivers
- A late-2026 Claude release could deliver a capability jump large enough to clear 55% on a single public leaderboard run.
- Anthropic can improve scores through better post-training, stronger reasoning scaffolds, and higher test-time compute without needing a brand-new paradigm.
- The market’s majority Yes pricing suggests traders believe the current trajectory is already close to the threshold.
- The outcome depends on the best public score in all of 2026, so Anthropic gets multiple chances to post a qualifying result.
Risk factors
- Humanity’s Last Exam is very hard, so incremental progress may be insufficient to bridge the last few percentage points.
- Anthropic may not optimize specifically for this benchmark, leaving the public score below the threshold even if the model improves broadly.
- A delayed release or slow leaderboard update could leave too little time for a qualifying Claude score to appear in 2026.
- Diminishing returns near the frontier can make each additional point more expensive and less predictable than earlier gains.
Scenarios
Best case
Anthropic releases a materially stronger Claude model in the second half of 2026, and the official leaderboard records a score comfortably above 55%, possibly in the high 50s or low 60s.
Most likely
Claude improves further during 2026 and gets near the threshold, with the final public score landing around the mid-50s and a meaningful chance of just clearing 55%.
Worst case
Claude improvements continue but plateau below the threshold, with the best public 2026 score ending in the high 40s or low 50s and never reaching 55%.
More from this day
- FinancialsKalshi13y
Will OpenAI or Anthropic IPO first?
AI24%MKT94%Edge-70HypedI think Anthropic is more likely to IPO before OpenAI, so if the market is asking whether OpenAI goes first, my answer is no more often than yes. Recent reporting has shifted the lead toward Anthropic, while OpenAI looks more likely to wait for a better valuation and a calmer market.
- CompaniesKalshi1y
Starbucks total global stores in 2026
AI68%MKT15%Edge+53Hidden GemStarbucks is more likely than not to finish 2026 above 41,800 global stores. The threshold is only modestly above its last known scale, and ordinary net expansion should get it there unless management materially slows growth or closes a large number of stores.
- PoliticsKalshi2y
Which agencies will Trump eliminate?
AI74%MKT30%Edge+44Hidden GemUSAID already looks functionally dismantled, with most staff, projects, and operational authority stripped away. The main uncertainty is whether the market resolves on de facto elimination or requires formal statutory abolition, but I still see a Yes outcome as more likely than not.