Quantifying Error Tolerance in Synthetic Data: An Atomic-level Operand vs. Operator Perturbation Study

📅 2026-08-29
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入Atomic Tree Operation Modeling (ATOM)框架,解决了合成数据生成中误差容忍度量化分析不足的问题,区分了操作数和操作符扰动,优化了数据过滤策略。
📝 Abstract
Synthetic data generation has become a cornerstone for advancing large language models. However, the lack of the quantitative analysis for error tolerance became a critical bottleneck. Consequently, current filtering strategies fluctuate between two extremes: they are either overly aggressive, risking the exclusion of potentially valuable samples, or overly permissive, failing to eliminate erroneous samples effectively. To bridge this gap, this paper introduces Atomic Tree Operation Modeling (ATOM), a framework that decomposes data into functional units ($f(x)\rightarrow y$). ATOM distinguishes benign Operand $x$ perturbations from fatal Operator $f$ perturbations. The former are needlessly discarded by aggressive filtering, while the latter slip through permissive filtering. Our experiments reveal a double dissociation: models are robust to operand perturbations but collapse under operator perturbations. By prioritizing operator over aggressive operand precision, our ATOM-synthesized data outperforms rigorous baselines (e.g., +3.1% gain over LIMA), suggesting that operator diversity matters more than operand precision. Our code is available at https://github.com/Lut-hub/ATOM.
Problem

Research questions and friction points this paper is trying to address.

error tolerance
synthetic data generation
filtering strategies
Innovation

Methods, ideas, or system contributions that make the work stand out.

Atomic Tree Operation Modeling (ATOM)
operand perturbations
operator perturbations
error tolerance
🔎 Similar Papers
No similar papers found.
Jiaxiang Liu
Jiaxiang Liu
Zhejiang University
Multimodal FusionMedical Image Analysis
C
Chenhao Yuan
C2DL, Institute of Automation, Chinese Academy of Sciences, Beijing, China
Shuwen Xu
Shuwen Xu
Xidian university, Mcmaster university, Full Professor
signal processing
B
Boxuan Xing
C2DL, Institute of Automation, Chinese Academy of Sciences, Beijing, China
X
Xiusheng Huang
C2DL, Institute of Automation, Chinese Academy of Sciences, Beijing, China
Y
Yinhao Xu
School of Artificial Intelligence, University of Chinese Academy of Sciences
H
Hao Liu
C2DL, Institute of Automation, Chinese Academy of Sciences, Beijing, China
W
Wenhao Teng
Department of Gastrointestinal Surgery Fujian Provincial Cancer Hospital
X
Xiangwen Liao
College of Computer and Data Science, Fuzhou University
Pengfei Cao
Pengfei Cao
Institute of Automation, Chinese Academy of Sciences
Natural Language ProcessingLarge Language ModelsInformation Extraction
J
Jun Zhao
C2DL, Institute of Automation, Chinese Academy of Sciences, Beijing, China
K
Kang Liu
C2DL, Institute of Automation, Chinese Academy of Sciences, Beijing, China