Which company has the third best AI model end of July?
I assess a modest but non-negligible 20% chance that Google will occupy the third spot on the Chatbot Arena Text Arena leaderboard on July 31, 2026, with the most likely outcome being that strong competitors push Google just outside the top three.
Analysis
The Chatbot Arena leaderboard is a community-driven, head-to-head preference ranking that can move rapidly as new models are released and as evaluation samples shift; it rewards conversational quality as judged against other models rather than raw benchmark numbers, so relative improvements or changes in sampling can produce sizable rank swings even without large underlying model changes. Because I have no live snapshot of the current leaderboard in your prompt, this assessment synthesizes known dynamics from 2023–2026: major vendor releases, community evaluation bias, and the frequency with which the top positions change following new model launches or gating/API changes.
Google has many structural advantages that increase the plausibility of landing third: deep engineering resources, large-scale training infrastructure, extensive multimodal research, and a history of producing competitive instruction-tuned models; if Google issues a timely model update between now and the end of July that is specifically optimized for conversational preference and Arena-style interactions, that could move it into the top three. However, Chatbot Arena rewards not just raw capability but specific alignment with human preference tests and sometimes punishes overzealous safety filtering or excessive refusal behaviors, which are real headwinds for large incumbent vendors focused on conservative deployments.
The competitive landscape is dense and fast-moving: OpenAI, Anthropic, Mistral, xAI, and community-tuned open models have all demonstrated the ability to leapfrog rankings when they introduce targeted improvements, pricing/access changes, or when the Arena’s evaluation set skews toward their strengths. Access and latency differences also matter because Arena evaluations are often run under conditions that favor responsive, broadly available models; if Google’s endpoint is rate-limited or tuned differently, that could depress Arena results independent of the underlying model’s raw capability.
Bringing these angles together, a probability in the low double digits reflects the combination of Google’s real capacity to reach a podium spot if they deliver a targeted update, versus the substantial and growing risk that nimble competitors or evaluation idiosyncrasies will occupy the top three by late July; this justifies a probability materially above a longshot but well under even odds, and slightly above the current market-implied 11% given Google’s resources and historical performance potential.
Arguments
For
- Google has the engineering and data resources to release a model capable of occupying a top-three leaderboard slot within a short timeframe.
- Google’s research and product pipeline historically produce substantial jumps in conversational capability when major updates are deployed.
- If Google tunes a model specifically to maximize human preference outcomes, it can outperform more generally optimized rivals in head-to-head Arena tests.
- Alphabet’s scale and investment in multimodal, instruction tuning, and safety infrastructure allows balancing capability with acceptable behavior that judges reward.
Against
- Other vendors and open-source communities have shown faster, more frequent targeted improvements that often change leaderboard order rapidly.
- Conservative safety tuning and refusal behaviors common in Google releases can lower Arena preference scores relative to more permissive models.
- Chatbot Arena’s evaluation methodology and sampling can favor models with particular stylistic tendencies that do not align with Google’s design choices.
- Access limitations, rate limits, or higher latency during evaluation periods can materially harm Google’s effective Arena ranking regardless of model quality.
Key drivers
- Timing and scope of any Google model update or instruction-tuning release before July 31.
- How well Google optimizes the model for conversational preference judged by Arena-style head-to-head evaluations.
- Competitor release cadence and improvements from OpenAI, Anthropic, Mistral, xAI, and community/tuned LLMs.
- Chatbot Arena sampling and evaluation set composition, which can advantage different model styles.
- Model access, latency, and rate limits during Arena evaluation windows, which affect real-world match outcomes.
Risk factors
- Google may prioritize safety and conservative refusals that reduce Arena-preference scores despite strong capabilities.
- Competitors could ship targeted updates or prompt-tuning that specifically improve Arena performance just before the cutoff.
- Leaderboard volatility and small-sample noise can produce rank movements that do not reflect broad capability differences.
- API gating, regional availability, or rate limits could exclude or degrade Google’s Arena performance at check time.
- Tiebreaker rules and exact unrounded scores could push Google down if scores are extremely close between rivals.
Scenarios
Best case
Google releases a focused conversational update or variant optimized for human-preference head-to-heads and enjoys strong, consistent Arena match wins, putting it solidly in third place or higher by July 31.
Most likely
Google remains highly competitive but narrowly misses the podium, finishing just outside the top three (roughly fourth to sixth) as one or more rapidly iterating competitors occupy the top positions on the Arena leaderboard.
Worst case
Competitors deliver targeted improvements and Google’s Arena endpoint is throttled or exhibits conservative refusal behavior, leaving Google well below third place and possibly outside the top five.
More from this day
- PoliticsKalshi2y
Which agencies will Trump eliminate?
AI99%MKT32%Edge+67Hidden GemUSAID has already been functionally eliminated as an independent agency (shuttered July 1, 2025 and folded into State); I assess a ~99% chance it will be counted as eliminated during Trump’s term.
- PoliticsKalshi18y
Which G7 leader will leave next?
AI35%MKT98%Edge-63HypedI assess ~35% chance the UK Prime Minister (Keir Starmer) will be the first G7 leader to leave office; the market at ~98% for that outcome looks severely overstated given scheduled election timing and relative stability of other leaders.
- PoliticsKalshi2y
Which Supreme Court justices will resign during Trump's term?
AI8%MKT63%Edge-55HypedIndependent assessment: very unlikely that Justice Samuel Alito will resign during Trump's 2025–2029 term; I estimate an 8% chance he voluntarily resigns during that window.