Institution profile

Zhengzhou University

Academic institutionasia · cn
Official website
Research library164linked papers
Opportunities0open roles
Selected work

Representative Papers

LangForce: Bayesian Decomposition of Vision Language Action Models via Latent Action Queries

Jan 21, 2026

This work addresses the limited generalization of current vision-language-action (VLA) models to novel instructions or multi-task settings, which stems from a collapse in the conditional mutual information between language instructions and actions—caused by the redundancy of instructions in training data where actions can be predicted directly from visual inputs alone. To tackle this “information collapse,” the paper formally characterizes the problem and introduces a Bayesian decomposition–based dual-branch architecture. This framework employs learnable latent action queries to separately model a vision-driven prior and a language-conditioned posterior, while explicitly optimizing pointwise mutual information (PMI) between actions and instructions to enforce instruction adherence. Evaluated on SimplerEnv and RoboCasa benchmarks, the method achieves significant generalization gains without additional data, improving out-of-distribution accuracy by 11.3% on SimplerEnv.

2 citationsRead paper
Recent publications

Latest Papers