TEFM: Token-Efficient Faithful Modeling for Structured Data

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出TEFM框架,通过压缩结构化数据为紧凑的行为代码来提高令牌效率,并采用双保真目标优化以保证忠实度,解决了LLM在关键领域应用中的令牌效率和忠实性问题。
📝 Abstract
In this paper, we solve two fundamental obstacles in applying LLMs to critical domains: token efficiency and faithfulness. To address both constraints jointly, we present TEFM (Token-Efficient Faithful Modeling), a framework designed for structured data analysis in critical domains. TEFM achieves token efficiency by compressing lengthy structured observations into compact Behavioral Code tokens, dramatically reducing token consumption with minimal information loss. Moreover, TEFM enables faithful rationalization through a dual-fidelity objective that jointly optimizes code-level reconstruction and prediction-level fidelity, identifying minimal sufficient feature subsets grounded in input data. Comprehensive experiments across various domain datasets and model backbones (Qwen3, Gemma-2, Phi-4) show that TEFM achieves competitive classification accuracy with dramatic token reduction (approximately 1\% token retention in clinical and 2\% in security domains) while producing faithful rationales.
Problem

Research questions and friction points this paper is trying to address.

token efficiency
faithfulness
structured data
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-Efficiency
Faithful Rationalization
Behavioral Code Tokens
Dual-Fidelity Objective
🔎 Similar Papers
No similar papers found.