Frequency-Conditioned Flow Matching for Vision-Language-Action Models

📅 2026-09-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了机器人动作频率异质性问题,通过引入FreqFM框架,该方法在DCT频率坐标中构建谱匹配源分布,并自适应平衡频率目标。
📝 Abstract
Robot actions are temporally correlated trajectories whose frequency components encode motion at different scales with highly non-uniform energy distributions. Yet Flow Matching--based vision-language-action (VLA) models typically generate actions in temporal coordinates, without explicitly modeling or systematically leveraging this frequency heterogeneity. We introduce \emph{FreqFM}, a frequency-conditioned Flow Matching framework for VLA models. It raises action frequency from an implicit trajectory property to an explicit conditioning dimension that spans the entire generation pipeline. Concretely, in DCT frequency coordinates, FreqFM constructs a spectrum-matched source distribution, adaptively balances the objective across frequencies, and constrains per-frequency guidance residuals using the corresponding reference transport scales. FreqFM integrates into existing Flow Matching action experts without changing the VLA backbone. Across LIBERO, LIBERO-Plus, and VLA-Arena, FreqFM consistently improves performance, including a 9.3-point gain on LIBERO-Plus, and further demonstrates its effectiveness on six real-robot tasks.
Problem

Research questions and friction points this paper is trying to address.

Flow Matching
vision-language-action models
frequency components
non-uniform energy distributions
Innovation

Methods, ideas, or system contributions that make the work stand out.

FreqFM
Frequency-Conditioned Flow Matching
DCT Frequency Coordinates
Spectrum-Matched Source Distribution
Adaptive Objective Balancing
🔎 Similar Papers
2024-06-09Annual Meeting of the Association for Computational LinguisticsCitations: 13
💼 Related Jobs
No related jobs found.
H
Haochen Niu
AGIBOT
S
Shengye Dong
AGIBOT, Xi'an Jiaotong University
H
Hao Liu
AGIBOT
P
Peiwen Lin
AGIBOT
W
Wang Chuang
AGIBOT