Highest Claude score on Humanity’s Last Exam in 2026?
Claude looks genuinely close to 55%, so a modest improvement or a new 2026 release could get the market to Yes. Still, the official evidence currently appears to sit just below the line, and the source inconsistency makes the higher market price look somewhat aggressive; I would put this around 69%.
Analysis
The central fact is that Claude is already very near the threshold. The most credible public figures in the current context place Anthropic's best official or near-official Humanity's Last Exam result around 53.3%, with another cited result at 52.6%, so only a small additional improvement is needed to reach 55%. That small gap matters because for frontier models, a one- to two-point gain on a hard benchmark can arrive with a new model revision, a better inference setup, or a more optimized evaluation run.
At the same time, the benchmark reporting landscape is messy, and that cuts both ways. Some trackers report higher Claude scores, including values above 55% when tools are allowed or under slightly different conditions, but the market resolves on the official Humanity's Last Exam leaderboard and on the HLE Accuracy measure specifically. If the public leaderboard is the only thing that counts, then a score that looks strong in an alternate setup may not help unless Anthropic posts a comparable official result that clearly clears the threshold.
The time remaining is enough for another attempt, but not enough to assume one. Anthropic has strong incentives to keep pushing frontier performance, and a year-end release could plausibly add the needed couple of points. Still, the market price implies a much higher chance than the raw official evidence alone suggests, so I think traders are leaning heavily on the possibility of a late-year leap or on already-available noncomparable results; my independent read is that Yes is favored, but not nearly as strongly as the current price indicates.
Arguments
For
- Arguments for Yes: Claude is already close enough that a small improvement could push it above 55%.
- Arguments for Yes: Anthropic has a strong incentive to keep improving benchmark performance before year end.
Against
- Arguments against Yes: The strongest clearly comparable official numbers still appear to be below 55%.
- Arguments against Yes: Reported scores above 55% may not match the exact official conditions used for resolution.
Key drivers
- Claude is already within a small margin of the 55% threshold, so a modest improvement would be enough.
- Anthropic still has several months left in 2026 to ship a stronger model or rerun the benchmark.
- Some reported Claude results already exceed 55% under different evaluation conditions, suggesting the ceiling is near.
- The official leaderboard may lag or differ from other trackers, which creates uncertainty about what will ultimately count.
Risk factors
- The best clearly comparable official score may remain stuck around the low 53% range.
- Any higher scores from other trackers may rely on tools or evaluation setups that do not match the resolution standard.
- If Anthropic does not release another major Claude upgrade in 2026, the threshold may not be crossed.
- Benchmark gains tend to get harder near the frontier, so the last couple of points are not guaranteed.
Scenarios
Best case
Anthropic releases or reruns a stronger Claude model on the official leaderboard, and the HLE Accuracy score lands in the mid-50s or higher before the end of 2026.
Most likely
Claude improves a bit further from its current level, and the final official result ends up either just above 55% or just below it, making the last update decisive.
Worst case
No official Claude result reaches 55% on the qualifying Humanity's Last Exam leaderboard, and the market resolves No despite some higher-looking alternative tracker results.
More from this day
- pop culturePolymarketEnded
"Spider-Man: Brand New Day" Opening Weekend Box Office (Higher Strikes)
AI84%MKT8%Edge+76Hidden GemMost credible domestic forecasts sit well below $280 million, so I think the under-threshold outcome is much more likely than the market price suggests. The main upside risk is that Spider-Man can still produce a record-level opening, but that would require a major surprise versus current tracking.
- pop culturePolymarketEnded
"Spider-Man: Brand New Day" Opening Day Box Office
AI73%MKT2%Edge+71Hidden GemThe current tracking suggests a very large opening, but not necessarily one large enough to clear 120m on the opening day figure used by this market. I think less than 120m is more likely than the market price implies, with the center of gravity in the low-to-mid 100s rather than comfortably above the line.
- pop culturePolymarketEnded
"The Odyssey" 3rd Weekend Box Office
AI71%MKT9%Edge+62Hidden GemThe latest box office tracking and industry forecasts point to a third weekend around the mid-40 millions, which puts the under-47m outcome in the lead. The market appears to be pricing in a meaningful chance of a stronger-than-expected hold, but the balance of evidence still favors Yes.