Reading is not Reasoning: Bridging the Agentic Policy Gap in Vision-Text Compression

πŸ“… 2026-08-09
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Although vision-to-text compression can reduce contextual costs in multi-step language agents, the modality shift induces a significant degradation in strategic capabilities, manifesting as policy drift in action selection, query formulation, termination judgment, and evidence utilization. To address this, this work proposes CAPS, the first framework to systematically uncover and analyze this cross-modal policy gap. CAPS introduces a two-stage cross-modal policy self-distillation mechanism that combines offline trajectories with online reinforcement learning, enabling knowledge transfer and capability preservation by distilling from the model’s textual-history policy to supervise its visual-history policy. Evaluated on SearchQA and ALFWorld, CAPS improves task success rates by up to 15.6% over AgentOCR while reducing average and peak memory context costs by up to 63.3% and 83.4%, respectively.
πŸ“ Abstract
Multi-step language-model agents repeatedly process growing interaction histories, leading to substantial context costs. Vision--text compression reduces these costs by rendering history as images, but the resulting modality shift creates a marked capability gap. Through controlled evaluations of history recovery, matched-state decisions, and complete trajectories, we show that this gap cannot be explained by OCR quality alone. Visual-history agents exhibit systematic drift in action selection, query formulation, stopping, and evidence use, revealing an agentic policy gap. We introduce \textbf{CAPS}, a two-stage \textbf{C}ross-modal \textbf{A}gentic \textbf{P}olicy \textbf{S}elf-distillation framework that uses the same model's stronger text-history policy to supervise its visual-history counterpart. Offline trajectory self-distillation transfers successful text-policy behavior to visual-history inputs, while online policy self-distillation provides dense supervision on states visited by the visual-history policy during reinforcement learning. On SearchQA, CAPS improves over AgentOCR by 5.0\% and 3.4\% with 3B and 7B backbones, respectively. On full-history ALFWorld, the corresponding gains are 15.6\% and 14.5\%. Across settings, CAPS reduces average memory-context cost by up to 63.3\% and peak cost by up to 83.4\% relative to matched text-history policies. These results show that explicit cross-modal policy self-distillation can preserve agent capability under vision--text compression. Our code will be made publicly available in a future release.
Problem

Research questions and friction points this paper is trying to address.

vision-text compression
agentic policy gap
context cost
multi-step agents
modality shift
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-text compression
agentic policy gap
cross-modal self-distillation
context efficiency
multimodal agents
πŸ”Ž Similar Papers
πŸ’Ό Related Jobs
No related jobs found.
Cheng Fan
Cheng Fan
Shenzhen University
Intelligent buildingsData miningBig dataBuilding energy conservation.
J
Junyi Zhou
City University of Hong Kong
T
Tingzhang Luo
City University of Hong Kong
R
RongJian Xu
City University of Hong Kong
Q
Qiyanhui Lu
City University of Hong Kong
M
Mingjian Zhu
Huawei Technologies Ltd.
Hanting Chen
Hanting Chen
Noah's Ark Lab, Huawei
deep learningmachine learningcomputer vision
Jianyuan Guo
Jianyuan Guo
City University of Hong Kong (CityU)