A Close Look At World Model Recovery In Supervised Fine-Tuned LLM Planners
This study investigates whether supervised fine-tuning enables large language models to acquire world-model representations and reasoning capabilities for end-to-end planning. To this end, we propose an interpretability framework tailored to planning-oriented large language models, integrating linear probing, internal representation analysis, and generative evaluation to systematically examine how models encode action validity and state predicates. Our experiments reveal that internal representations can linearly separate valid from invalid actions—even when the output layer fails to classify them accurately—and demonstrate that the breadth of state-space coverage in the training data significantly influences the fidelity of recovered world models, underscoring the critical role of data diversity in shaping planning competence.