Highest Meta score on Humanity’s Last Exam in 2026?
Meta has a plausible path to a 55% HLE score in 2026, but I think the bar is still a bit high given the difficulty of the benchmark and the limited time left this year. My estimate is that No is slightly more likely than Yes.
Analysis
The core question is whether any Meta-published model in 2026 can reach at least 55% accuracy on Humanity’s Last Exam before the year ends. That is a demanding threshold on a benchmark built to separate strong general models from truly exceptional ones, so the challenge is not simply releasing a bigger model but producing a system that improves enough on difficult reasoning, knowledge retrieval, and robustness to clear a fairly high line. With only a few months left in 2026, the market is effectively asking whether Meta already has a near-threshold model in development or can compress a meaningful leap into a short remaining window.
Arguments for Yes are real because Meta has both the resources and the flexibility to try multiple approaches before the deadline. The market only needs the best Meta model published during 2026 to reach the threshold, so an earlier model can fail while a later one succeeds. If Meta ships a larger foundation model, a stronger reasoning-focused variant, or a release with more aggressive post-training and test-time compute, the score could improve nonlinearly rather than gradually. Hard benchmarks also sometimes show sharp gains once a lab aligns training, inference strategy, and data curation well enough, so a late-year release could still surprise to the upside.
Arguments against Yes are stronger to me because 55% feels like a meaningful jump rather than a marginal improvement, and Meta has not historically been the most aggressive benchmark-maximizer on the hardest frontier evaluations. Even if Meta continues improving its models, the remaining time is short relative to the engineering, evaluation, and deployment cycle needed to deliver a system that is clearly above the cutoff. In addition, these very hard exams often resist easy scaling gains, so routine iteration may leave Meta in the high 40s or low 50s without quite crossing the line. The current market price near fifty-fifty reflects the uncertainty, but my read is that the threshold is just hard enough that No deserves a modest edge.
Arguments
For
- Arguments for Yes: Meta can still release one or more improved 2026 models, and only the best one needs to clear 55%.
- Arguments for Yes: Better post-training, tool use, or test-time compute could produce a sharp score jump rather than a linear one.
Against
- Arguments against Yes: Meta has often lagged the very top labs on the hardest reasoning-style benchmarks.
- Arguments against Yes: Reaching 55% before year-end likely requires a substantial step up, not just routine model iteration.
Key drivers
- Meta still has multiple chances in 2026 to publish a stronger model before the deadline.
- A single late-year reasoning-focused release could create a nonlinear jump in HLE accuracy.
- The 55% threshold is high enough that small incremental gains may not be sufficient.
Risk factors
- Meta may prioritize broad product quality and efficiency over pure leaderboard chasing.
- Humanity’s Last Exam is designed to be resistant to easy benchmark gains.
- A breakthrough model could appear late, but the remaining timeline is short enough to make execution risk meaningful.
Scenarios
Best case
Meta ships a major late-2026 model with much stronger reasoning and inference-time strategy, and its best HLE score clearly exceeds 55% before the deadline.
Most likely
Meta releases at least one better model in 2026, but the top score finishes close to the cutoff and ends up a bit under 55%, making No the outcome.
Worst case
Meta’s 2026 models improve but stall below the threshold, with the best score ending in the high 40s or low 50s and No resolving.
More from this day
- politicsPolymarket3mo
Google Maps renames Lake Ontario to "Lake America" by...?
AI5%MKT95%Edge-90HypedThis looks overwhelmingly unlikely. Google Maps would need to roll out a broad, production-level U.S. renaming of a major international lake within months, and there is no sign of that happening.
- pop culturePolymarketEnded
"Spider-Man: Brand New Day" 5th Weekend Box Office
AI90%MKT3%Edge+87Hidden GemThe under-20m outcome looks strongly favored for a Spider-Man film in its fifth weekend. Unless it is holding far better than typical superhero releases, the domestic weekend should be comfortably below the threshold.
- pop culturePolymarketEnded
"The Dog Stars" Opening Weekend Box Office
AI41%MKT93%Edge-52HypedThe under on 8 million is plausible, but the broader tracking picture still leans a bit above that line. I think the market is slightly overpricing the chance of a sub-8 million opening, so I put Yes at 41%.