Highest Claude score on Humanity’s Last Exam in 2026?
The balance of evidence strongly favors Yes because multiple recent reports already place at least one Claude model at or above 55% on Humanity’s Last Exam in 2026. The main uncertainty is not capability but whether the official resolution source preserves the same score definition and leaderboard snapshot.
Analysis
The central point is that the reported 2026 Claude results are already at the threshold or above it in several recent summaries, with one Claude model listed at 55% and another at 56% on Humanity’s Last Exam. Since the market only needs the highest Claude-family score in 2026 to reach 55% or more, those figures, if reflected on the official leaderboard, are enough to make Yes the natural outcome. This is a materially stronger position than a typical benchmark market where the target has not yet been reached, because the threshold appears to have been crossed already rather than merely approached.
The biggest reason not to assign absolute certainty is that the recent evidence is not perfectly consistent. Different summaries appear to refer to different snapshots, and at least one source gives a lower number for a Claude model or blends in a tools-enabled variant that may not map cleanly to the exact official HLE Accuracy the market will use. In other words, the risk is less about whether Claude can achieve 55% and more about whether the official source and this market’s definition of the metric line up cleanly enough for that score to count. If the leaderboard later shows a lower canonical value, or if the highest score on the official page is reinterpreted as a different variant, the market could still surprise to No despite strong outside evidence.
From a market-pricing perspective, the current Yes price around 87.75% looks somewhat conservative relative to the evidence that the threshold has already been met. That said, prediction markets sometimes leave room for resolution friction, and this event has exactly the sort of wording where ambiguity around metric naming, benchmark variants, or source availability can matter. Because of that, the market should not be treated as a 99% lock, but the base rate and the available public reporting both point heavily toward Yes unless the official leaderboard behaves very differently from the recent summaries.
Arguments
For
- Arguments for Yes: Multiple recent summaries already show Claude at or above the 55% threshold.
- Arguments for Yes: The market only needs the best Claude score in 2026 to clear 55%, which appears to have happened already.
Against
- Arguments against Yes: The evidence is messy and could reflect different benchmark variants rather than the exact official HLE Accuracy.
- Arguments against Yes: If the official leaderboard later shows a lower canonical score or a withdrawn result, the market could resolve No.
Key drivers
- Recent 2026 reports already place at least one Claude model at 55% or higher on Humanity’s Last Exam.
- The market resolves from the official leaderboard, so the decisive issue is whether that source confirms the reported score.
- The market threshold is low relative to the reported Claude results, making a Yes outcome the default if the figures are accurate.
Risk factors
- The publicly reported scores are not fully consistent, which creates a chance that the official metric is being summarized differently across sources.
- If the official leaderboard uses a different HLE variant, snapshot, or naming convention, a reported 55% may not count the way traders expect.
Scenarios
Best case
The official Humanity’s Last Exam leaderboard clearly shows a Claude model at 55% or 56% on the exact HLE Accuracy metric, confirming Yes with no ambiguity.
Most likely
The official leaderboard ultimately matches the recent reports closely enough that a Claude model is recognized at or above 55%, so the market resolves Yes.
Worst case
The reported scores turn out to be from a non-canonical variant or tool-enabled setting, and the official 2026 leaderboard never shows a Claude model at 55% on the resolved metric.
More from this day
- pop culturePolymarketEnded
What will be the top global Netflix movie this week?
AI62%MKT10%Edge+52Hidden GemThe Whisper Man looks meaningfully favored to finish as the week’s top global Netflix movie because it is already showing up as the leading title in recent Netflix movie roundups and has strong launch-week visibility. Still, the outcome is not close to certain because the official weekly ranking can shift if another title sustains better full-week viewing.
- pop culturePolymarketEnded
What will be the top US Netflix movie this week?
AI19%MKT68%Edge-49HypedI think The Whisper Man has a meaningful but limited shot at reaching number one, and the market looks a bit too optimistic at 25%. My estimate is 19% Yes because a Netflix title can spike quickly, but it still needs unusually strong broad appeal to beat the field.
- EconomicsKalshi9y
US real GDP growth in 2035?
AI63%MKT15%Edge+48Hidden GemMy independent estimate puts the most likely 2035 U.S. real GDP growth outcome in the 1.6% to 2.5% range, with the center of gravity around the BLS-style 2.1% long-run pace. The market appears somewhat too concentrated on very low-growth or near-zero outcomes relative to the available long-run projections.