Which company has the best AI Agent end of September?
Anthropic is still a leading contender and may well finish first, but the market appears to be pricing in too much certainty given how quickly the leaderboard can move. I think Yes is more likely than not, but far from a lock.
Analysis
Anthropic enters the final stretch of September with a real advantage in the current public narrative. The freshest composite model rankings have placed Claude Fable 5.1 at or near the top, and Anthropic’s recent product expansion around computer access gives it a plausible edge in a market that rewards practical agent behavior, not just raw chat quality. If the leaderboard reflects broad model reputation, overall agent usefulness, and sustained performance across several task types, Anthropic has a strong chance to remain the company in first place at the check time.
That said, the key weakness in the Yes case is that the field is not standing still, and the most relevant benchmark evidence is mixed. OpenAI’s GPT-6 Astra appears to be ahead on explicit browser-agent and autonomous web-navigation tasks, which are especially important if the leaderboard’s practical evaluation weights direct action-taking more heavily than general model polish. A model can look strongest in one composite ranking while still losing narrowly on a task suite that better matches real agent work. That makes the 95 percent market price look stretched, because the underlying evidence supports Anthropic as a favorite, not as a near-certain winner.
The broader competitive context also argues against overconfidence. There are multiple active players improving agent products quickly, and the market has less than a month for a lead to change with a single model refresh or a leaderboard methodology quirk. Because the resolution depends on one specific snapshot at one specific time, even a small reversal, tie, or ordering nuance could matter. Anthropic is well positioned, but the combination of active competition, measurement noise, and the possibility that browser-style performance matters more than reputation suggests a materially lower probability than the current market implies.
My baseline is that Anthropic remains the most likely single company to hold first place at the end of September, but the edge is not large enough to justify near-certainty. The strongest interpretation of the evidence favors Anthropic in a close race, while the strongest counterinterpretation says OpenAI has already shown better evidence on the most agent-relevant benchmarks. That balance supports a solid Yes probability, but one that still leaves substantial room for a No outcome.
Arguments
For
- Arguments for Yes: Anthropic currently has credible evidence of broad model-quality leadership and strong general agent positioning.
- Arguments for Yes: Its recent product expansion in computer access could help sustain a top leaderboard rank through the resolution date.
Against
- Arguments against Yes: OpenAI appears stronger on browser-agent and autonomous task benchmarks that may better reflect true agent performance.
- Arguments against Yes: The leaderboard is volatile enough that a small late move or methodology-sensitive shift could remove Anthropic from first place.
Key drivers
- Anthropic has recently held the top spot in at least one composite model ranking, which supports its current lead.
- OpenAI’s newest agent benchmark strength creates a credible path for Anthropic to lose first place before month-end.
Risk factors
- A late-September model update from OpenAI could overtake Anthropic on the leaderboard snapshot.
- Small ranking differences or tie-breaking rules could flip the winner even if the two companies remain very close.
Scenarios
Best case
Anthropic’s current lead in composite rankings holds through the September 30 snapshot, and no rival model overtakes Claude Fable 5.1 or another Anthropic model at the top of the Agent Arena list.
Most likely
Anthropic remains one of the top two contenders and may well finish first, but the margin stays close enough that late leaderboard movement or a task-specific advantage from OpenAI keeps the outcome genuinely uncertain.
Worst case
OpenAI’s GPT-6 Astra or another competitor improves enough on agent tasks to take first place by the check time, pushing Anthropic into second or lower.
More from this day
- CompaniesKalshi1y
Starbucks total global stores in 2026
AI64%MKT11%Edge+53Hidden GemStarbucks has a credible path to clear 41,800 global stores by its 2026 reporting date, and I think the market is pricing in too little growth. My independent estimate is a 64% chance of Yes, not 11%.
- FinancialsKalshi13y
Will OpenAI or Anthropic IPO first?
AI41%MKT91%Edge-50HypedI think Anthropic is more likely to IPO before OpenAI, so the Yes outcome is only moderately likely. The current market looks too optimistic on OpenAI getting to the finish line first given the newer guidance that OpenAI is likely a 2027 story while Anthropic is being discussed in near-term IPO windows.
- PoliticsKalshi2y
Who will Trump pardon?
AI4%MKT50%Edge-46HypedBarron Trump receiving a pardon before 2029 looks very unlikely based on the available information. The main reason is simple: there is no sign he needs one, and a pardon requires an underlying offense or at least a credible legal predicament to resolve.