Short answer: among general-purpose AI models, OpenAI's reasoning models hold the best chess record. In the only sanctioned event ever held - Google's Kaggle Game Arena chess exhibition, August 2025 - OpenAI's o3 won the title and swept the final 4–0 against xAI's Grok 4, with Google's Gemini taking the podium spot behind them. That's the entire official record of AI-vs-AI chess, and OpenAI owns it.
The only sanctioned podium in existence
| Place | Model | Lab | Note |
|---|---|---|---|
| 🥇 1 | o3 | OpenAI | Swept the final 4–0 |
| 🥈 2 | Grok 4 | xAI | Furthest run by a non-OpenAI fighter |
| 🥉 3 | Gemini 2.5 Pro | Google DeepMind | Podium at its own company's event |
Kaggle Game Arena chess exhibition, August 2025 - the only sanctioned LLM chess event to date.
The three asterisks that matter
- Chess engines still crush every LLM. Stockfish would beat any model on this page without warming up. The interesting question was never "can AI play chess" (solved decades ago) - it's whether general intelligence can, without a search engine bolted on.
- Most LLM chess losses are illegal moves, not checkmates. Community leaderboards consistently show chat models forfeiting by playing moves that don't exist. Legality-under-pressure is the real skill gap.
- The weirdest data point in AI: an older model, gpt-3.5-turbo-instruct, plays roughly 1800-Elo chess while many newer chat models flail - strong evidence that chess skill in LLMs is a training-data quirk, not general capability. Full story: why can't LLMs play chess?
One exhibition is a snapshot, not a season. VERSUZ runs chess as a rated, scheduled event from September 1, 2026 - same fighters, long series, engine-refereed, with a live per-game Elo that replaces every claim on this page with a scoreboard. The preseason matchups are on the fight card.