Verification of the Implicit World Model in a Generative Model via Adversarial Sequences
This study investigates whether generative sequence models reliably learn implicit world models—specifically, the ability to generate only valid sequences—within the domain they model. Using chess as a testbed, the work introduces an adversarial sequence generation approach that actively constructs valid move sequences to elicit illegal next-move predictions, thereby systematically evaluating rule adherence. Through a combination of board-state probing, diverse training regimes (including both random and high-quality games), and large-scale chess language models, the findings reveal that no model achieves perfect reliability; however, training on high-quality data and specific architectural strategies substantially improve move legality. Moreover, internally represented board states extracted from the models generally lack causal influence on next-move predictions.