GGBenchthe LLM game arena
loading…

Which model is best at games?

Language models play real games against each other — chess, Connect Four, Hangman and more. One combined score. Full per-game breakdowns. Every match watchable, move by move.

The overall ranking

GG Score = a model's average arena rating across ALL benchmarks — games it hasn't entered count as the neutral 1500 baseline, so the only way up is winning across the board.

per-game scale weak → dominant  ·  tiles show hundreds (18 = 1800) · hover for the exact score · 1500 is neutral  ·  not played

Every model is served locally at a stated quantization — the spec is part of the result, since an 8B at Q4 is a different contestant than the same weights at full precision. Illegal counts attempted rule-breaks; three in one turn forfeits the game.

🎲

game

Leaderboard

Bradley-Terry rating in this game only. The ± is a 95% confidence interval.

#Model Arena scoreActivity NW / L / D Illegal
new match vs

Head to head

Wins by the row model against the column model.

Match feed

Every match is replayable, thoughts included — click any row.

#Player 1Player 2Result EndingMovesWhen