CHSM8-player-base: games-only continuation (one full pass over 26M games)
CHSM8 pretrained decoder (LegumMagister/chsm8-pt) continued for exactly one pass over 26.0M unique human games (1.95B move tokens; Lichess pool, deduplicated, split seed 20260802) with no text: plain next-move prediction on the whole game (both sides). Unconditioned strong base for the player fine-tunes (see Player models).
- Architecture: Llama-style decoder, 12 layers x 960, 15 heads, SwiGLU 3520, QK-norm, factored move head; 166.2M parameters (chsm8-pt without its 44.3M cross-attention scaffold).
- Pretraining cross-attention (frozen, constant input) removed exactly: its constant output became a per-layer bias (
xbias); loss at step 0 identical to the pretrained model. - Recipe (picked by 10/30/60-min ablations, val loss): AdamW lr 5e-5, cosine to 0 by tokens, warmup 5%, batch 4 GPUs x 32 x 512, legal-move penalty 0.5, bf16.
- Val loss (fixed 40-batch val set) at 10%, 20%, β¦ 90% of the pass: 2.9236, 2.9117, 2.9104, 2.9068, 2.8948, 2.8871, 2.8855, 2.8786, 2.8773
Strength: 1000 games vs Stockfish 17.1 (UCI_Elo 2000, 0.05 s/move)
| checkpoint (share of the pass) | Elo |
|---|---|
| pct20 | 1889 (95% CI 1869-1907, 999 games) |
| pct40 | 1881 (95% CI 1861-1900, 1000 games) |
| pct60 | 1886 (95% CI 1867-1905, 1000 games) |
| pct80 | 1883 (95% CI 1863-1902, 999 games) |
| final | 1890 (95% CI 1870-1909, 1000 games) |
Reference: the pretrained base, chsm8-pt, scored 1837 on the same benchmark; one extra pass over games gives +53.
Player models
Each player model is this model fine-tuned on one player's own moves (full fine-tune, AdamW lr 3e-5, cosine, 2 passes over the player's unique games; the opponent's moves are context only). Games come from chess.com and lichess for active players and from PGN Mentor (classical over-the-board) for historical ones. Held out: the player's most recent games (val + test, up to 1,000 each), never trained on.
How well a fine-tune captures the player (on held-out test games, player's own moves only):
- Move-match: share of positions where the model's top move is the move the player actually played (base = this model before fine-tuning).
- Top-3: the played move is among the model's top 3.
- CE: cross-entropy on the player's moves (lower = closer to the player).
- Model Elo: the fine-tuned model's strength, 1000 games vs Stockfish 17.1 at UCI_Elo 2000, 0.05 s/move (this base model: 1890).
The model imitates the player's moves in the games available (mostly online blitz and rapid for active players); it does not reach the player's strength. Fine-tunes stay within about Β±40 Elo of the base.
FIDE top 10
| # | Player | FIDE (2026-09) | Model | Unique games | Move-match base β fit | Top-3 | CE base β fit | Model Elo |
|---|---|---|---|---|---|---|---|---|
| 1 | Magnus Carlsen | 2823 | chsm8-player-carlsen | 17,561 | 0.554 β 0.569 | 0.798 | 2.969 β 2.915 | 1903 |
| 2 | Hikaru Nakamura | 2792 | chsm8-player-nakamura | 19,997 | 0.540 β 0.572 | 0.797 | 3.039 β 2.862 | 1850 |
| 3 | Fabiano Caruana | 2784 | chsm8-player-caruana | 5,513 | 0.545 β 0.556 | 0.788 | 2.960 β 2.938 | 1902 |
| 4 | Javokhir Sindarov | 2782 | chsm8-player-sindarov | 1,371 | 0.557 β 0.567 | 0.794 | 2.959 β 2.897 | 1897 |
| 5 | Nodirbek Abdusattorov | 2780 | chsm8-player-abdusattorov | 5,356 | 0.544 β 0.549 | 0.782 | 2.967 β 2.995 | 1906 |
| 6 | Wesley So | 2770 | chsm8-player-so | 5,300 | 0.547 β 0.564 | 0.793 | 2.962 β 2.969 | 1910 |
| 7 | Praggnanandhaa R | 2763 | chsm8-player-praggnanandhaa | 7,467 | 0.543 β 0.550 | 0.776 | 3.016 β 3.008 | 1871 |
| 8 | Vincent Keymer | 2761 | chsm8-player-keymer | 10,906 | 0.538 β 0.567 | 0.789 | 3.073 β 2.980 | 1887 |
| 9 | Wei Yi | 2758 | chsm8-player-wei-yi | 2,869 | 0.552 β 0.555 | 0.780 | 3.007 β 3.015 | 1895 |
| 10 | Alireza Firouzja | 2757 | chsm8-player-firouzja | 19,999 | 0.555 β 0.559 | 0.785 | 2.962 β 2.960 | 1884 |
Bonus: historical players
Four classic players, added because they shaped chess history and are widely known. Their games are classical over-the-board games from before online chess, so there is less data and much older play.
| Player | Peak FIDE | Why | Model | Unique games | Move-match base β fit | CE base β fit | Model Elo |
|---|---|---|---|---|---|---|---|
| Garry Kasparov | 2851 (1999) | 13th World Champion; world #1 for most of 1984β2005 | chsm8-player-kasparov | 1,727 | 0.533 β 0.539 | 3.153 β 3.103 | 1900 |
| Bobby Fischer | 2785 (1972) | 11th World Champion; won the 1972 "Match of the Century" against Spassky | chsm8-player-fischer | 660 | 0.572 β 0.567 | 2.969 β 2.891 | 1875 |
| Anatoly Karpov | 2780 (1994) | 12th World Champion (1975β85); one of the most successful tournament players ever | chsm8-player-karpov | 3,128 | 0.519 β 0.543 | 3.159 β 3.043 | 1905 |
| Judit Polgar | 2735 (2005) | strongest woman player in history; world #8 overall at her peak | chsm8-player-polgar-judit | 1,459 | 0.530 β 0.539 | 3.181 β 3.099 | 1887 |
Small corpora (Fischer: 660 games) give noisier fits: Fischer's top-1 move-match drops slightly while his CE improves.
Usage
The checkpoint is a PyTorch dict (model_state_dict, cfg) for the CHSM8 decoder (factored move head over
(kind, src, dst, piece, promo)). Moves are fed as factored rows, BOS first; there is no text input, and play picks the
best legal move by summed factored log-probability. The model classes and the loading and play code will be released later with the project code.