Publication
Player-Specific Policy Adaptation in Elite Chess with MoE-LoRA and Limited Search
Authors:
Loris Sogliuzzo, Aloïs Rautureau, Eric Piette
Venue:
38th Benelux Conference on Artificial Intelligence (BNAIC 2026), 2026 (accepted)
Topics:
Human-AI Alignment, Behavioral Modeling, Policy Adaptation, Mixture of Experts, Monte Carlo Tree Search, Jensen-Shannon Divergence
Abstract
The historical trajectory of chess artificial intelligence has predominantly focused on superhuman playing strength, yielding policy models whose decision paths optimize purely for game-theoretic value rather than fidelity to human behavior. While recent human-centric neural architectures successfully model population-wide behavior across Elo rating intervals, personalizing policies to capture the individual styles of elite players and World Champions remains an open challenge. Furthermore, standard evaluations rely on deterministic top-1 move accuracy, which fails to capture the natural stochasticity and strategic variance inherent in human play.
This paper investigates player-specific behavioral modeling across a curated cohort of 14 elite chess players whose international careers began in the 20th century, encompassing six World Champions and eight World Championship contenders or Candidates-level elite players. Because fine-tuning player embeddings over a frozen backbone fails to alter decision boundaries effectively, a parameter-efficient Mixture of Experts with Low-Rank Adaptation (MoE-LoRA) framework is introduced, applying dynamically gated residual adjustments directly to the network logit space. To reduce the risk of loops and mitigate tactical blunders without collapsing human stylistic variance, this architecture is coupled with budget-constrained Monte Carlo Tree Search regularized by top-p Nucleus Pruning. Behavioral alignment is evaluated strictly on recurring board positions with verified empirical depth (n ≥ 5 observations per player) to eliminate degenerate singleton Dirac collapses. Experimental results show that this framework improves top-1 move accuracy from 43.61% to 63.02%, reduces the tactical gap (ΔCPL) from 10.49 to 1.47, and minimizes state-averaged Jensen-Shannon Divergence to 0.2368 when regularized by search.
Full reference
Sogliuzzo, L., Rautureau, A., Piette, E. (2026). Player-Specific Policy Adaptation in Elite Chess with MoE-LoRA and Limited Search. 38th Benelux Conference on Artificial Intelligence (BNAIC 2026). Accepted.
BibTeX
@inproceedings{sogliuzzo2026playerSpecific,
author = {Sogliuzzo, Loris and Rautureau, Alo{\"i}s and Piette, Eric},
title = {Player-Specific Policy Adaptation in Elite Chess with {MoE-LoRA} and Limited Search},
booktitle = {38th Benelux Conference on Artificial Intelligence (BNAIC 2026)},
year = {2026},
note = {Accepted}
}