
maxim-saplin/llm_chess on GitHub
- Stars
- 131
- Forks
- 15
- Language
- HTML
- License
- Apache-2.0
- Last push
- 22 Sep 2026
Read from GitHub on . Counts change daily.
LLM Chess makes language models play chess in an agentic loop, against either a random-move player or the Komodo Dragon engine, and scores two different things at once: how well the model actually plays (win/loss record, Elo) and whether it can sustain the interaction without breaking the protocol, for example by producing an illegal or hallucinated move that ends the game early, tracked separately as game duration.
The project’s own findings, from testing a range of models over time, are that 2024-era models often struggled just to follow the move format, while stronger 2025 reasoning models started winning almost every game against the random player, which is why Dragon was added as a tougher opponent to keep the ranking meaningful. Jev is among the models the project has run through this harness, alongside models from OpenAI, Anthropic, Google and others configured through its .env setup, plus local models via Ollama or LM Studio. The leaderboard data lists jev-latest, added on 2026-09-17, with 80 games (8 wins, 50 losses, 22 draws), no illegal moves, every game played to the end, and an estimated Elo of about 243 (plus or minus 118).
Results live on a public leaderboard, and the underlying paper was presented at NeurIPS’s FoRLM workshop in 2025. The benchmark is open source and runs against your own API keys.

