Highest Google Gemini score on Humanity’s Last Exam in 2026?
Gemini has been improving quickly on related reasoning benchmarks, but the only directly cited Humanity’s Last Exam result is still 37.5% without tools, which leaves a meaningful gap to 50%. I think the market is overestimating the chance of a threshold-crossing score by year-end, though the probability is still material.
Analysis
The most important hard data point is the latest cited Gemini result of 37.5% on Humanity’s Last Exam without tools. That is a strong score in absolute terms, but it is still 12.5 percentage points below the market threshold, and closing that gap on a difficult benchmark in only a few remaining months is not trivial. Since the market resolves on the highest Gemini score in 2026, not just the current score, the key question is whether Google can ship a substantially better model before the end of the year rather than merely continue incremental improvement.
There are real arguments that a jump is possible. Google’s Gemini family is showing very strong performance on related expert-level reasoning tests, which signals that the underlying model stack is improving and that another generation or a better inference setup could produce a noticeably higher HLE score. If Google releases a new flagship model, uses more compute at inference, or leans on a leaderboard-eligible tool-augmented configuration, the score could move enough to cross 50. The main limitation is that HLE is a tougher and less forgiving benchmark than many adjacent tests, so strong progress elsewhere does not automatically translate into a qualifying result here.
The current market price implies a fairly optimistic view of a near-term breakthrough, but the concrete evidence does not yet support that level of confidence. The latest cited Gemini HLE score is still below the threshold, while the fact that other non-Google systems have cleared 50% shows the target is achievable in principle but not that Gemini is likely to do it on this timeline. With only a few months left in 2026, I see a genuine chance of a late-year leap, but not a majority likelihood. My assessment is that No is more likely than Yes, even though the Yes case cannot be dismissed.
Arguments
For
- Arguments for Yes: Gemini is already strong on closely related reasoning benchmarks, so another substantial improvement is plausible.
- Arguments for Yes: There is still enough time in 2026 for a new release or better inference stack to add the points needed to clear 50%.
Against
- Arguments against Yes: The best directly cited Gemini HLE score is only 37.5% without tools, which is far from the 50% threshold.
- Arguments against Yes: HLE appears harder to scale than adjacent benchmarks, so strong performance elsewhere may not translate into a qualifying score.
Key drivers
- Whether Google launches a meaningfully stronger Gemini model before the end of 2026.
- Whether official leaderboard scoring allows a tool-augmented or otherwise advantaged Gemini setup to count.
- How much HLE performance can improve from the current 37.5% baseline within one product cycle.
Risk factors
- Google could release a late-2026 Gemini variant that makes a larger-than-expected leap on HLE.
- The leaderboard could favor an inference or tool configuration that boosts Gemini more than the current score suggests.
- Rapid benchmark progress across frontier models could make a 12.5-point improvement more feasible than historical patterns imply.
Scenarios
Best case
Google ships a much stronger Gemini model or a leaderboard-eligible tool-assisted setup that pushes the official HLE score to 50% or higher before year-end.
Most likely
Google continues to improve Gemini but does not make a large enough leap on Humanity’s Last Exam to reach 50% in 2026.
Worst case
Gemini improves only modestly and tops out in the low-to-mid 40s, leaving the best 2026 score below the threshold.
More from this day
- PoliticsKalshi2y
Which agencies will Trump eliminate?
AI66%MKT23%Edge+43Hidden GemUSAID has a materially better-than-even chance of being functionally eliminated or dismantled during Trump’s term, even if the exact legal termination path is messy. My independent estimate is well above the market’s 23% Yes price because the political direction, staffing pressure, and prior Trump-era hostility all point toward a serious elimination attempt.
- pop culturePolymarketEnded
# of views of Grand Theft Auto VI Extended Look on week 1?
AI40%MKT79%Edge-39HypedThe market looks too bearish on the view count. Grand Theft Auto VI content is one of the few gaming videos that can plausibly clear 20 million views in a week even without a major celebrity or event tie-in, so I lean toward No.
- techPolymarket3mo
Next Mythos-Class Model: Text Arena Debut?
AI71%MKT94%Edge-23HypedAnthropic’s next Mythos-class model looks more likely than not to clear 1470 if it appears publicly on the Arena leaderboard, but the market is pricing in too much confidence because visibility and leaderboard qualification are still meaningful hurdles. I would put the chance of a Yes at 71%.