Will a Chinese company have a top ___ AI model by December 31?
I put the chance of a Chinese company reaching the top 3 on the Arena text leaderboard by year-end at 16%. Chinese labs are improving quickly, but the current top tier is still dominated by Anthropic and the best Chinese model is still well outside the top 3.
Analysis
As of July 10, the overall Text Arena top 3 is entirely Anthropic, with claude-fable-5 at 1509, claude-opus-4-6-thinking at 1504, and claude-opus-4-7-thinking at 1503. The highest-ranked Chinese model visible on the overall board is qwen3.7-max-preview at rank 17 with 1475, while GLM-5.2 max, DeepSeek, Xiaomi, Baidu, and ByteDance all sit in the next cluster behind it. That means the current Chinese frontier is competitive, but still meaningfully below the line needed to displace any of the current top 3.
The main argument for a Yes outcome is that Chinese labs are still iterating fast and keep showing up on the arena with new releases. Arena's own changelog shows new Chinese entries being added to the text leaderboard, including Qwen3-235b-a22b-instruct-2507, and Qwen's recent Arena post said Qwen3.7 Preview had landed on Arena and pushed Alibaba to a top-six text lab position. Z.ai's GLM-5.2 launch also emphasizes long-horizon work and a 1M-token context window, and Reuters-based coverage says GLM-5.2 has already narrowed the gap with top US systems.
The main reason to stay cautious is that the bar is not just high, it is currently occupied by multiple strong frontier systems from the same competitor, Anthropic. Because this market only needs one checked moment, a short-lived breakout would count, but a Chinese model still has to beat several entrenched frontier systems before December 31. Reuters also reported possible limits on overseas access to China's most advanced models, which could reduce visibility on Arena if those plans are implemented. On balance, that leaves me above the current market price, but still well below coin-flip territory.
Arguments
For
- Chinese labs are shipping new models quickly, which keeps the door open for a late-year breakout.
- GLM-5.2 is being positioned as a flagship long-horizon model with a 1M-token context window, a capability leap that could translate into a stronger arena score.
- Qwen3.7 Preview recently landed on Arena and Alibaba is already close enough to the front pack that one good revision could matter.
- The market checks any time before year-end, so even a brief top-3 appearance would satisfy the condition.
Against
- Anthropic currently occupies the top three overall slots, so Chinese models must clear a large and visible gap.
- The best Chinese model on the overall board is still only around rank 17, which shows the current frontier gap is real.
- Several Chinese models are competitive but still clustered below the leaders rather than breaking into the absolute top tier.
- If Chinese firms restrict overseas access to their strongest models, Arena may see less exposure and fewer votes from the global user base.
Key drivers
- The most important driver is whether any Chinese lab can deliver a score jump of roughly 25 to 35 points from current levels.
- Release cadence matters because this market can resolve on a temporary peak, not just a year-end standing.
- Z.ai, Alibaba, Moonshot, DeepSeek, and ByteDance all have credible paths to another major release before year-end.
- The current frontier benchmark gap is smaller than it was a year ago, so a single strong model generation could plausibly reshuffle the upper ranks.
Risk factors
- The current top three are already very strong and would need to be surpassed, not merely matched.
- Chinese models may continue to improve without quite reaching the extreme quality jump needed for the top 3.
- Arena scores can move around after public launch, so a promising preliminary result may not hold long enough.
- Any policy or access restrictions that limit global testing could make it harder for Chinese models to gather the votes needed for a top-3 showing.
Scenarios
Best case
A Chinese flagship model from Z.ai, Alibaba, Moonshot, or ByteDance lands in the top 3 during a release window, and the market resolves Yes even if the stay is brief.
Most likely
Chinese models remain very competitive and may continue climbing, but the top 3 stays dominated by US labs through December 31.
Worst case
Anthropic, OpenAI, and Google keep the front line locked down, Chinese releases improve but stay a few dozen Elo points short, and no Chinese model ever checks into the top 3.
More from this day
- PoliticsKalshi3mo
Will a cabinet member be impeached?
AI97%MKT5%Edge+92Hidden Gem**Very likely Yes.** A sitting Cabinet member (Philippine Vice President Sara Duterte) is already being tried in the Senate impeachment court, which in the Philippines follows a House impeachment — so the factual threshold for “impeached” has effectively been met and the probability of at least one Cabinet member being impeached before 2027 is extremely high.
- techPolymarketEnded
Best AI model on July 18?
AI13%MKT97%Edge-84HypedI think claude-opus-4-6-thinking is a clear underdog here. It is currently near the top of the leaderboard, but it is not the leader today, and the latest frontier releases make it more likely to stay around second or third than to reclaim first.
- CompaniesKalshi1y
Robinhood funded customers in 2026
AI12%MKT81%Edge-69HypedBased on reported mid-2026 metrics and required growth rates, I assess a low probability that Robinhood will report above 30.2M funded customers in 2026 — roughly a 12% chance.