Skip to content
System One

maxim-saplin/llm_chess

A benchmark that has language models play chess against a random-move player and the Komodo Dragon engine, scoring both chess skill and whether the model keeps making legal moves without hallucinating.

Image from maxim-saplin/llm_chess
Image from github.com

maxim-saplin/llm_chess on GitHub

Stars
131
Forks
15
Language
HTML
License
Apache-2.0
Last push
22 Sep 2026

Read from GitHub on . Counts change daily.

LLM Chess makes language models play chess in an agentic loop, against either a random-move player or the Komodo Dragon engine, and scores two different things at once: how well the model actually plays (win/loss record, Elo) and whether it can sustain the interaction without breaking the protocol, for example by producing an illegal or hallucinated move that ends the game early, tracked separately as game duration.

The project’s own findings, from testing a range of models over time, are that 2024-era models often struggled just to follow the move format, while stronger 2025 reasoning models started winning almost every game against the random player, which is why Dragon was added as a tougher opponent to keep the ranking meaningful. Jev is among the models the project has run through this harness, alongside models from OpenAI, Anthropic, Google and others configured through its .env setup, plus local models via Ollama or LM Studio. The leaderboard data lists jev-latest, added on 2026-09-17, with 80 games (8 wins, 50 losses, 22 draws), no illegal moves, every game played to the end, and an estimated Elo of about 243 (plus or minus 118).

Results live on a public leaderboard, and the underlying paper was presented at NeurIPS’s FoRLM workshop in 2025. The benchmark is open source and runs against your own API keys.

Open on GitHub

More like this

Keep browsing