How Output Format Confounds Data Quality and Capability in Instruction Tuning

📅 2026-09-01
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过12项任务和多种模型家族,揭示了输出格式对数据质量和模型能力评估的影响,并提出使用梯度特征等方法解决这一混淆问题。
📝 Abstract
Instruction-tuning data are judged by quality metrics, and tuned models are judged by benchmarks, but both judgments pass through an output interface: the surface format in which an answer is written. Using gradient signatures across 12 tasks, four semantically equivalent interfaces, three model families, and controlled corruptions, we show that this interface confounds both measurements. Spectral statistics such as effective rank are provably invariant to interface rotation and empirically blind to semantic corruption, while the direction of the update carries the quality signal. The interface-varying residual is not noise: it identifies each unit's own target task perfectly across all three families. Capability itself is stored relative to the training interface: a skill that raises accuracy by more than 40 points under the training format can be nearly invisible under every other, and correcting a single generation budget flips the measured effect of fine-tuning on GSM8K from a gain into a large loss. Pre-registered interventions delimit where this geometry stops short of control. Data quality and model capability are interface-conditioned quantities, and current practice often reports the interface instead of the content.
Problem

Research questions and friction points this paper is trying to address.

Output Format
Data Quality
Model Capability
Instruction Tuning
Innovation

Methods, ideas, or system contributions that make the work stand out.

gradient signatures
interface confounding
effective rank
semantic corruption
model capability
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
C
Chengguang Gan
Independent Researcher
H
Hanjun Wei
University of Chinese Academy of Sciences
Y
Yunhao Liang
University of Chinese Academy of Sciences
Qinghao Zhang
Qinghao Zhang
Department of Electrical Engineering, Tsinghua University
ReliabilityPower electronicsjunction temperaturecondition monitoringlifetime prediction
S
Shiwen Ni
Shenzhen University of Advanced Technology
Zhixi Cai
Zhixi Cai
Research Fellow, Monash University
computer visiondeepfakemultimodalvisual reasoningllm agent