Reconciling Contradictory Views on the Effectiveness of SFT in LLMs: An Interaction Perspective
This study investigates why supervised fine-tuning (SFT) is consistently effective for small models yet yields inconsistent or even detrimental results in large language models (LLMs). To address this, the work introduces a novel perspective by analyzing the evolution of inter-token interactions during SFT, leveraging interaction-based interpretability techniques to quantify and track dynamic changes in interaction strength. The findings reveal that in LLMs, SFT primarily acts as a denoising mechanism within an extremely short training window, after which it rapidly overfits, leading to performance degradation. This insight offers a new understanding of the boundaries of SFT effectiveness and empirically validates the necessity of early stopping across multiple LLMs and datasets, providing practical guidance for optimizing fine-tuning protocols.