Train What You Deploy:Token-Faithful Post-Training of a Production Coding

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出了一种保真度感知的训练框架和C-DPPO方法,解决了编码和终端代理后训练中的令牌与控制保真度错误问题。
📝 Abstract
Existing post-training pipelines for coding and terminal agents suffer severe token and control fidelity errors: simplified training environments mismatch production deployments, and offline token reconstruction from agent logs distorts original prompts and conflates policy calls with background model operations. We present a fidelity-aware training coupling framework that retains trainer-side sampling over original prompts, eliminates spurious model calls via a negotiated training protocol, and restricts loss computation to verifiable token spans with closed-failure guarantees. We further propose Certified Divergence Proximal Policy Optimization (C-DPPO), which establishes tight two-sided TV certification bounds, adaptive-K rules, budget-aware sequence guarantees, and error-robust policy masking atop standard DPPO. Evaluated on matched Baize5B and Baize10B models with identical training and test protocols on TMax-100, C-DPPO yields a consistent +3.0-point performance gain over standard DPPO across model scales. Certificate audits validate the reliability and full operational coverage of our certified training pipeline.
Problem

Research questions and friction points this paper is trying to address.

post-training
token fidelity
control fidelity
Innovation

Methods, ideas, or system contributions that make the work stand out.

fidelity-aware training
Certified Divergence Proximal Policy Optimization (C-DPPO)
token-faithful post-training
🔎 Similar Papers
No similar papers found.