Which company has best AI model end of July?
I assess a low but non-negligible chance (~20%) that Google will occupy the top spot on Chatbot Arena's LLM leaderboard on July 31, 2026, because Google has the resources to leapfrog competitors but faces short-term disadvantages from competitor momentum, evaluation biases, and timing uncertainty.
Analysis
Market-implied probability (Yes ~10.7%) shows traders view Google as an underdog for the July 31 Chatbot Arena top rank, likely reflecting recent leaderboard trends favoring OpenAI/Anthropic-style models, the small time window remaining, and the possibility that Google will not launch a clearly superior public chat model in the next three weeks. The marketplace has substantial volume, meaning this pricing aggregates many informed views, but it may underweight tail events like a surprise Google release or an aggressive model deployment specifically tuned to Arena preferences.
Historically, Chatbot Arena rankings reflect human pairwise preference tests and have tended to reward conversational quality, safety tradeoffs, and accessible interactive behavior rather than raw capability benchmarks alone, which has benefited companies that prioritize user-facing chat optimization and rapid iteration on conversational style. Google is a top-tier R&D organization with multimodal assets and the capacity to deliver large quality improvements, but past public appearances of Google chat models have sometimes emphasized safety controls and conservatism in responses that can depress preference scores compared with more freely responsive competitors.
Key uncertainties over the next three weeks include whether Google will release a new or updated model that is accessible and configured in a way that performs well in Arena-style human preferences, whether competitors (notably OpenAI and Anthropic) will push new updates or targeted tweaks before July 31, and whether the Arena evaluation environment or sampling could change; any of these could shift rankings rapidly. Given these dynamics, I place modest probability on Google overtaking the field: it is plausible as a technical and product matter, but the combination of competitor momentum, the short timeline, and evaluation idiosyncrasies means the most likely outcome remains that another company holds first place on the check date.
Arguments
For
- Google has world-class research teams, infrastructure, and data that can produce top-tier LLMs if they prioritize a rapid release cycle.
- If Google deploys a new model specifically tuned for chat and human preference signals, it could leapfrog competitors on Arena-style evaluations.
- Google’s multimodal and retrieval-augmented capabilities could yield conversational responses that users prefer for many real-world prompts.
- A surprise early release or a public demo optimized for Arena could generate a short-term advantage in pairwise human evaluations.
- Google can marshal extensive testing and fine-tuning resources to iterate quickly on conversational trade-offs if the goal is to top leaderboards.
Against
- Recent historical patterns on human-preference leaderboards have favored models from OpenAI and Anthropic that prioritize conversational responsiveness.
- Google often emphasizes safety and conservative behaviors which can reduce preference scores in direct human pairwise comparisons.
- The remaining timeframe is short for a transformative public release, especially if internal testing or legal/regulatory reviews delay deployment.
- Competitors are highly motivated and capable of rapid incremental improvements that may close or widen Google’s gap before July 31.
- Chatbot Arena results can be sensitive to prompt distributions and evaluator samples that may not reflect Google’s strengths.
- Alphabetical tie-breaker rules could work against Google if scores end up extremely close and Google is tied with a company earlier in the alphabet.
Key drivers
- Timing and scale of any Google model release or update before July 31 that is public and present on Chatbot Arena.
- How Google configures safety, helpfulness, and verbosity trade-offs that directly affect pairwise human preferences in Arena tests.
- Competitors' update cadence—OpenAI, Anthropic, xAI, and others may push incremental or major improvements before the check.
- The Chatbot Arena sampling population, evaluation prompts, and any temporary changes to evaluation protocols that favor certain conversational styles.
- Google's ability to tune and rapidly iterate on user-facing chat behavior to maximize human-preference win rates.
- Model accessibility and integration: models must be available on the Arena leaderboard in the required format to be eligible for ranking.
Risk factors
- Google may choose not to push a public model update tuned for Arena-style human preferences in the short window before July 31.
- Competitors could release stronger models or targeted patches specifically optimized for human pairwise comparisons, locking Google out of first place.
- The leaderboard resolution could be affected by unavailability or changes in Arena’s display or measurement practices, creating ambiguity or delay.
- Arena’s evaluation sample and prompt set could systematically favor conversational styles or risk tolerances that are not aligned with Google's design choices.
- Last-minute tie-breaker dynamics (identical rank and score) will revert to alphabetical ordering among tied companies, which could help or hurt Google depending on ties.
- Public access limitations or rate-limiting could prevent Google models from being effectively represented in Arena matchups at scale.
Scenarios
Best case
Google rolls out a new public chat model or a major update to an existing model before July 31, the model is included and well-represented on Chatbot Arena, human evaluators strongly prefer its conversational style and capabilities, and Google tops the leaderboard decisively due to superior responses on the Arena prompt sample.
Most likely
Google does not overtake the existing leader by July 31 because competitors retain momentum and the short time window limits the impact of any Google update, but Google remains competitive and could be near the top without securing first place.
Worst case
No public Google update appears or their release emphasizes conservatism and safety to the detriment of perceived helpfulness, competitors maintain or extend their lead through targeted improvements, and Google finishes well below first place on the Arena leaderboard.
More from this day
- PoliticsKalshi2y
Which Supreme Court justices will resign during Trump's term?
AI8%MKT66%Edge-58HypedIndependent assessment: very unlikely that Justice Samuel Alito will resign during Trump's 2025–2029 term; I estimate an 8% chance he voluntarily resigns during that window.
- pop culturePolymarketEnded
What will Trump say during Press Conference in Turkey?
AI60%MKT3%Edge+57Hidden GemI assess a moderately high probability that Trump will say the words "Million," "Billion," or "Trillion" 20+ times during the Turkey press conference, mainly because of his documented habit of repeating numeric magnitudes and the likelihood he will pivot to domestic economic talking points during Q&A.
- FinancialsKalshi13y
Will OpenAI or Anthropic IPO first?
AI35%MKT82%Edge-47Hyped**Assumption:** 'Yes' = OpenAI will IPO before Anthropic. Independent assessment: I assign a 35% probability that OpenAI will IPO before Anthropic (i.e., Anthropic is more likely to be the first of the two to list).