Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the interference caused by coupled visual and coding capabilities in chart-to-code generation by proposing the MoCA framework. The method decouples dual branches via a cross-modal arbitration block, employing a lightweight arbiter for token-level dynamic weight allocation, combined with self-distillation warm-up and reinforcement learning optimization. Experiments demonstrate that MoCA achieves competitive performance across three benchmarks. Ablation studies confirm that these gains stem from complementary initialization and conditional arbitration mechanisms rather than mere model scaling, effectively enabling precise synergy in structured generation capabilities.
📝 Abstract
Chart-to-code generation requires a model to read the fine-grained visual details of a chart and write executable code that reproduces it. Existing chart-to-code methods either train visual and coding abilities separately, or fine-tune on chart-to-code data with the two abilities entangled. Neither strategy accounts for the distinct nature of the two abilities or the interference that arises when they are optimized together. We propose MoCA (Mixture of Cross-modal Arbitration), which separates the two abilities rather than blending them. MoCA is built on Cross-modal Arbitration Block (CAB), which maintains a visual branch and a code branch as two distinct pathways, and a lightweight arbiter that arbitrates their relative contributions at every layer and generated token. We train MoCA in two stages: a supervised warm-up on self-distilled reasoning trajectories that decomposes visual understanding into explicit steps, followed by reinforcement learning with rewards on both the reasoning process and the final code. Analysis shows that the arbiter learns structured rather than arbitrary allocations, with expert contributions varying systematically across tokens, layers, and instances. Across three benchmarks, MoCA delivers competitive performance against general-domain and chart-specialized models. Ablation results show that the gains cannot be attributed to a larger model size alone, but instead arise from the joint contributions of complementary visual and code branch initialization and input-conditioned arbitration through CAB.
Problem

Research questions and friction points this paper is trying to address.

Chart-to-Code Generation
Cross-modal Interference
Visual Understanding
Code Generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Token-Level Modality Arbitration
Mixture of Cross-modal Arbitration
Cross-modal Arbitration Block
Chart-to-Code Generation
Reinforcement Learning
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Q
Qinghao Fu
Zhejiang University, Ant Group
Y
Yarong Wang
Zhejiang University, Ant Group
S
Shunlei Ning
Ant Group
Y
Yilin Wang
Zhejiang University, Ant Group
S
Shunwen Bai
Zhejiang University, Ant Group
Xinda Wang
Xinda Wang
University of Texas at Dallas
Software SecurityAI SecuritySystems Security
J
Jiaotuan Wang
Ant Group
Y
Yinan Nie
Fudan University
W
Wei Zhou
Ant Group