Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
研究解决了高质量预训练检查点在后续训练中表现不佳的问题,通过分析30B混合专家模型训练流程中的检查点质量,发现高解决方案密度的检查点更优。
研究解决了高质量预训练检查点在后续训练中表现不佳的问题,通过分析30B混合专家模型训练流程中的检查点质量,发现高解决方案密度的检查点更优。
This work addresses the susceptibility of conditional log-probability–based scoring in multiple-choice evaluation to answer length bias, which can unfairly penalize or overfavor longer responses. To mitigate this issue, the paper proposes Bayesian Accuracy—a plug-and-play scoring method grounded in Bayesian posterior probabilities. By explicitly modeling the prior distribution over answer lengths, the approach effectively eliminates linear length bias without requiring additional forward passes. Evaluated across diverse benchmarks and few-shot settings, Bayesian Accuracy substantially reduces empirical length bias and enhances the fairness and reliability of model assessment compared to standard accuracy and length-normalization baselines.
Existing game-theoretic models in adversarial domains such as law often overlook language as a mechanism of persuasion, failing to capture the nuanced, discourse-driven nature of strategic interaction. This work proposes a Strategic Courtroom Framework that treats language as a first-class strategic action space. It introduces a heterogeneous multi-agent system grounded in nine interpretable personality traits and incorporates a reinforcement learning–driven dynamic trait orchestrator to generate adaptive persuasive strategies tailored to opponents and case specifics. Evaluated using DeepSeek-R1 and Gemini 2.5 Pro across 10 synthetic cases, 84 three-trait combinations, and over 7,000 simulated trials, the framework demonstrates that heterogeneous agent teams significantly outperform homogeneous ones, and dynamically composed traits surpass handcrafted strategies, with quantitative and charismatic traits contributing most prominently to persuasive efficacy.
This work addresses the lack of systematic investigation into format selection and performance trade-offs in existing low-bit quantization-aware training (QAT) methods, as well as their insufficient evaluation on generative tasks. To this end, we propose the first integration of k-means clustering into QAT for 1-bit weight quantization, optimizing generative performance under a fixed inference memory budget. Our approach transcends the limitations of conventional integer-based quantization schemes by leveraging learned cluster centroids to better preserve model fidelity at ultra-low bitwidths. Experimental results demonstrate that, under identical memory constraints, our method significantly outperforms state-of-the-art integer quantization approaches while maintaining compatibility with general-purpose hardware for efficient deployment.
This work addresses the challenge of inefficient computational resource allocation at byte-sequence boundaries in end-to-end hierarchical sequence modeling. The authors propose a boundary enrichment metric \( B \) to quantify the alignment between chunk starting positions and regions of high prediction difficulty, and introduce the Sombrero method, which leverages a confidence-aligned boundary loss and input-level confidence-weighted smoothing to steer boundaries toward harder-to-predict segments. Notably, Sombrero achieves improved computational efficiency without requiring explicit chunking or reliance on specialized routers. Evaluated on a 1B-scale UTF-8 corpus encompassing English–German text, code, and mathematical content, the approach significantly enhances the trade-off between accuracy and efficiency by concentrating computational resources on the most challenging prediction locations.
研究解决了高质量预训练检查点在后续训练中表现不佳的问题,通过分析30B混合专家模型训练流程中的检查点质量,发现高解决方案密度的检查点更优。
This work addresses the susceptibility of conditional log-probability–based scoring in multiple-choice evaluation to answer length bias, which can unfairly penalize or overfavor longer responses. To mitigate this issue, the paper proposes Bayesian Accuracy—a plug-and-play scoring method grounded in Bayesian posterior probabilities. By explicitly modeling the prior distribution over answer lengths, the approach effectively eliminates linear length bias without requiring additional forward passes. Evaluated across diverse benchmarks and few-shot settings, Bayesian Accuracy substantially reduces empirical length bias and enhances the fairness and reliability of model assessment compared to standard accuracy and length-normalization baselines.
Existing game-theoretic models in adversarial domains such as law often overlook language as a mechanism of persuasion, failing to capture the nuanced, discourse-driven nature of strategic interaction. This work proposes a Strategic Courtroom Framework that treats language as a first-class strategic action space. It introduces a heterogeneous multi-agent system grounded in nine interpretable personality traits and incorporates a reinforcement learning–driven dynamic trait orchestrator to generate adaptive persuasive strategies tailored to opponents and case specifics. Evaluated using DeepSeek-R1 and Gemini 2.5 Pro across 10 synthetic cases, 84 three-trait combinations, and over 7,000 simulated trials, the framework demonstrates that heterogeneous agent teams significantly outperform homogeneous ones, and dynamically composed traits surpass handcrafted strategies, with quantitative and charismatic traits contributing most prominently to persuasive efficacy.
This work addresses the lack of systematic investigation into format selection and performance trade-offs in existing low-bit quantization-aware training (QAT) methods, as well as their insufficient evaluation on generative tasks. To this end, we propose the first integration of k-means clustering into QAT for 1-bit weight quantization, optimizing generative performance under a fixed inference memory budget. Our approach transcends the limitations of conventional integer-based quantization schemes by leveraging learned cluster centroids to better preserve model fidelity at ultra-low bitwidths. Experimental results demonstrate that, under identical memory constraints, our method significantly outperforms state-of-the-art integer quantization approaches while maintaining compatibility with general-purpose hardware for efficient deployment.
This work addresses the challenge of inefficient computational resource allocation at byte-sequence boundaries in end-to-end hierarchical sequence modeling. The authors propose a boundary enrichment metric \( B \) to quantify the alignment between chunk starting positions and regions of high prediction difficulty, and introduce the Sombrero method, which leverages a confidence-aligned boundary loss and input-level confidence-weighted smoothing to steer boundaries toward harder-to-predict segments. Notably, Sombrero achieves improved computational efficiency without requiring explicit chunking or reliance on specialized routers. Evaluated on a 1B-scale UTF-8 corpus encompassing English–German text, code, and mathematical content, the approach significantly enhances the trade-off between accuracy and efficiency by concentrating computational resources on the most challenging prediction locations.