Which company has the third best AI model end of July?
I assess a low probability that Google will occupy exactly the third slot on the Chatbot Arena Text Arena overall leaderboard on July 31, 2026, assigning a 20% chance based on platform participation uncertainty, strong competition, and short time to move positions.
Analysis
The market price (Yes ~9%) implies participants currently believe it is unlikely Google will be exactly third on the Chatbot Arena Text Arena leaderboard at the check time; that price likely reflects liquidity and traders hedging around the small possible outcome space of a single rank. I put higher odds than the market price because Google historically fields very competitive models and large-scale improvements or tiered rollouts could push its model into the top three, but still view those odds as modest given competing dynamics.
Historically (up to my last broad knowledge cutoff) Google has had world-class models and large engineering investment, giving it a realistic chance to be among the top-ranked models when it participates openly; however, being in the top three specifically is fragile because the ranking is zero-sum — a single strong release from OpenAI, Anthropic, Mistral or other challengers can change placement quickly. The Chatbot Arena ranking is based on crowdsourced pairwise comparisons and an Arena score rather than purely automated benchmarks, so differences in safety filtering, conversational style, and availability to evaluators can materially shift rankings even if quantitative model capabilities are similar.
Platform and logistics factors matter disproportionately for this event: only models accessible to Chatbot Arena evaluators and included in the Text Arena overall leaderboard count, and the market uses the table with style control off, which affects how some systems perform relative to their marketed capabilities. If Google’s current or new model is not present on the Arena or is accessible only in a mode disallowed by the leaderboard setting, its probability of landing third is essentially zero regardless of raw capability. The short remaining horizon (about 25 days) reduces the chance that a large, reliable ranking shift will occur unless a major, immediately accessible release appears on the Arena.
Competitive dynamics and sampling noise also weigh against a high probability: OpenAI and Anthropic (and increasingly strong open-source or startup entrants) have shown rapid improvements and are likely to remain near the top, compressing the distribution of scores beneath them and making exact ranking positions volatile; the Arena’s human evaluation sampling and tie-breaking rules (including underlying granular scores and alphabetical tiebreaker) add stochasticity that increases the chance Google ends up adjacent to third rather than exactly third. Taken together — platform participation uncertainty, fierce competition, short timeframe, and evaluation noise — I assess a 20% chance that Google will be exactly third on July 31, 2026.
Arguments
For
- Google has historically built high-performing large language models and can plausibly produce a model competitive with the top three.
- Google’s engineering resources and user data could enable an incremental or step change release that improves Arena pairwise outcomes quickly.
- If Google’s model is on the Arena and tuned for the prompts and style-control-off setting, it may score strongly with human evaluators.
- Alphabet’s willingness to iterate and ship model improvements could produce a late push that elevates Google into the third slot before the check.
Against
- The Chatbot Arena ranking rewards direct human preference in pairwise matches, and other vendors (OpenAI, Anthropic, Mistral) have repeatedly dominated those comparisons.
- If Google’s model is not accessible or intentionally limited on the Arena, it cannot occupy third regardless of underlying performance.
- Small differences in safety behavior and refusal rates can hurt human preference scores even when raw capability is similar.
- Rapid competitor releases or optimizations in July could displace Google or keep it outside the top three.
- Evaluation sampling variance and the auction-like nature of pairwise tests make exact positions noisy and sensitive to short-term swings.
- Alphabet may choose to gate features or model variants that reduce its Arena competitiveness relative to more permissive models.
Key drivers
- Whether Google’s newest conversational model is present and fully accessible on the Chatbot Arena Text Arena leaderboard at the resolution time.
- A near-term Google model improvement or public release that demonstrably improves human pairwise preferences on arena-style prompts.
- Competitor releases from OpenAI, Anthropic, Mistral, xAI, or other strong entrants that either push Google down or leave the top rankings unchanged.
- How safety filters, system prompts, and the 'style control off' leaderboard setting affect perceived conversational quality for Google relative to others.
- Sampling noise and evaluator population composition on Chatbot Arena during the period that determines the Arena score.
Risk factors
- Google’s model is not listed or is restricted on Chatbot Arena at the check time, making third place impossible regardless of capability.
- A major competitor release in July that jumps ahead of Google and compresses the leaderboard ranks.
- Arena evaluation methodology or prompt sets favor certain conversational styles that do not align with Google’s strengths.
- The short time window before July 31 leaves little runway for Google to make a decisive, globally observable improvement on the Arena.
- Leaderboard tie-breaking rules and granular score rounding could place Google just above or below third due to tiny margins.
Scenarios
Best case
Google publishes or makes accessible a clearly improved conversational model before July 31, the model is included on Chatbot Arena in an unrestricted mode, and human pairwise judgments place it solidly in third on the Text Arena overall leaderboard.
Most likely
Google is present on the leaderboard and remains competitive but finishes outside exactly third (most likely landing in a nearby slot such as fourth or fifth due to strong competitors and Arena sampling noise).
Worst case
Google’s model is absent or heavily restricted on the Arena, or multiple competitors release stronger accessible models in July, resulting in Google being outside the top three and the market resolving to No.
More from this day
- PoliticsKalshi3mo
Will a cabinet member be impeached?
AI99%MKT6%Edge+93Hidden GemBased on the reported May 11, 2026 House impeachment of Vice President Sara Duterte and the scheduled Senate trial (July 6, 2026), the factual condition for a 'Yes' has already occurred under the event's plain wording; I assess a 99% independent probability that the event will resolve Yes.
- pop culturePolymarketEnded
"Spider-Man: Brand New Day" Opening Weekend Box Office
AI85%MKT4%Edge+81Hidden GemI assess a high probability that Spider-Man: Brand New Day will open for less than $200M domestically on its opening weekend, with my best estimate at 85% chance of falling below that threshold based on franchise history, box office norms, and release-risk factors.
- PoliticsKalshi3mo
Will Trump invoke the Insurrection Act?
AI99%MKT19%Edge+80Hidden GemIndependent assessment: overwhelmingly likely (already occurred); I assign a 99% probability that Trump has invoked the Insurrection Act during his presidency based on multiple corroborating facts and public statements.