Which company has second best AI model end of July?
I assign Google a 28% chance of occupying second place on the Chatbot Arena Text Arena | Overall leaderboard at the July 31, 2026 12:00 PM ET check, higher than the current market price but reflecting substantial competition and leaderboard volatility.
Analysis
There is no fresh news feed available for this assessment, so I rely on known product trajectories and historical leaderboard behavior: Google’s Gemini family has been a consistent top-tier entrant in multi-turn conversational evaluations and receives frequent updates, which gives Google an ongoing chance to land in the top three on community evaluation leaderboards like Chatbot Arena. The market-implied probability (Yes 6.3%) suggests traders see this outcome as unlikely, but that price can reflect risk aversion and a bias toward incumbents like OpenAI or Anthropic rather than a model-by-model technical assessment.
Historically, Chatbot Arena rankings are volatile over short windows because they are driven by pairwise human comparisons, newly submitted models, and sampling noise; Google has frequently appeared near the top but not always in the exact same ordinal position across leaderboard snapshots, meaning a credible path exists to second place but it is far from guaranteed. Competitive dynamics matter: OpenAI and Anthropic have released iterations that often occupy the very top positions, while newer entrants (e.g., Mistral variants, xAI, or specialized fine-tuned models) can surge quickly on this kind of crowd-sourced evaluation, compressing the probability that any single company secures exactly the second slot.
Leaderboard-specific technicalities also shape probabilities: the resolution rule orders by rank, then exact Arena score (including granular decimals), and finally company alphabetical order as a tie-breaker, so small numeric score differences and how models are labeled can flip second place at the snapshot moment; sampling timing (the July 31 noon ET check) matters if a model release or leaderboard submission occurs shortly before that time. Taking those factors together, Google has meaningful structural advantages (scale of data, R&D cadence, product deployment experience) that justify a probability materially above the market’s 6%, but the crowded competitive field and the inherent leaderboard sampling noise keep the probability well below even odds, leading me to 28% as a balanced estimate.
Arguments
For
- Google has a track record of producing top-tier conversational models that routinely occupy high leaderboard positions.
- Frequent updates and large-scale infrastructure allow Google to iterate and improve performance ahead of a snapshot.
- Google’s training data scale and retrieval/knowledge integration can boost text-chat accuracy in human evaluations.
- If competitors focus on multimodal or niche strengths, Google’s general-purpose chat tuning may excel on the Text Arena metric.
- Alphabetical tie-breaker works in Google’s favor against companies listed later in the alphabet.
- Corporate resources and engineering depth reduce the chance of last-minute deployment failures compared with smaller teams.
Against
- OpenAI and Anthropic have often claimed the very top slots, making it more likely Google is first or third rather than exactly second.
- Newer entrants and specialized fine-tuned models can spike quickly in crowd evaluations and displace incumbents.
- Chatbot Arena’s pairwise human judgments produce significant short-term rank volatility that can erase expected advantages.
- If Google’s improvements are more multimodal-focused, pure text-chat performance on the Arena leaderboard might lag.
- Tie-breaker alphabetization could disadvantage Google if ties occur with companies earlier in the market’s listed order.
- Market noise and small-sample effects at the snapshot moment mean a single anomalous test set or rater pool can flip rank outcomes.
Key drivers
- Google’s ongoing Gemini model updates and heavy R&D investment which tend to push its conversational models into the top tier.
- The rapid pace of new model releases and fine-tuned entrants which can reorder rankings quickly in a short window.
- Chatbot Arena’s human pairwise-evaluation methodology that produces volatility and sampling noise in rank positions.
- Timing of any Google model release, benchmark submission, or a competitor release within days of the July 31 snapshot.
- Model selection and configuration used by Chatbot Arena for the Text Arena | Overall ranking, which may favor certain instruction tuning styles.
- Potential differences between multimodal strengths and pure text-chat performance that could advantage or disadvantage Google in a text-focused leaderboard.
Risk factors
- A major release or upgrade from OpenAI, Anthropic, Mistral, or another competitor just before July 31 that pushes Google down.
- Chatbot Arena sample composition changes or curator updates that systematically favor newer or open-source models.
- Score ties resolved alphabetically could disadvantage Google if tied with companies earlier in the alphabet and listed differently in the market group.
- Unexpected bugs, access limits, or throttling on Google’s model in the Arena that depress human evaluations at snapshot time.
- Community preference shifts in evaluation criteria (e.g., factuality vs. creativity) that favor other architectures.
- Low volume and market mispricing which can cause trader prices to diverge far from realistic technical probabilities.
Scenarios
Best case
Google ships a targeted Gemini update or leaderboard submission timed before the snapshot that demonstrably improves conversational quality on text-only tasks, while competitors either stagnate or their releases miss the snapshot window, resulting in Google cleanly taking second place with a clear Arena score gap.
Most likely
Google remains among the top handful of companies but leaderboard noise and active competition produce a close cluster of scores around first-to-fourth places, with Google landing either first, third, or fourth and only a modest chance of occupying exactly second at the precise snapshot time, yielding roughly a one-in-four probability of second place.
Worst case
A rival release from OpenAI, Anthropic, or a rapidly improving open-source model gains strong human-evaluation momentum right before the snapshot, and either Google falls to third-or-worse or a small score tie plus alphabetical ordering pushes Google off second place entirely.
More from this day
- PoliticsKalshi2y
Which Supreme Court justices will resign during Trump's term?
AI8%MKT66%Edge-58HypedIndependent assessment: very unlikely that Justice Samuel Alito will resign during Trump's 2025–2029 term; I estimate an 8% chance he voluntarily resigns during that window.
- pop culturePolymarketEnded
What will Trump say during Press Conference in Turkey?
AI60%MKT3%Edge+57Hidden GemI assess a moderately high probability that Trump will say the words "Million," "Billion," or "Trillion" 20+ times during the Turkey press conference, mainly because of his documented habit of repeating numeric magnitudes and the likelihood he will pivot to domestic economic talking points during Q&A.
- FinancialsKalshi13y
Will OpenAI or Anthropic IPO first?
AI35%MKT82%Edge-47Hyped**Assumption:** 'Yes' = OpenAI will IPO before Anthropic. Independent assessment: I assign a 35% probability that OpenAI will IPO before Anthropic (i.e., Anthropic is more likely to be the first of the two to list).