MOSS-VL Technical Report

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the absence of concurrent perception-generation capabilities and response latency in real-time visual-language model interactions by treating real-time interaction as a first-class capability through full-stack co-design. We propose a gated cross-attention mechanism, synthetic interactive data supervision, and a lightweight fine-tuning curriculum to jointly optimize streaming perception and generation. Experiments demonstrate that the proposed model achieves state-of-the-art performance across multiple streaming benchmarks, with significantly superior proactive alerting capabilities. Notably, its advantage in time-to-first-token latency widens as visual context length increases, effectively enhancing the overall real-time interactive experience.
📝 Abstract
We present MOSS-VL, an open vision-language model family that treats real-time interaction -- perceiving while it speaks -- as a first-class capability. It is co-designed across the stack: the language decoder attends to vision only through gated cross-attention, so the model can naturally see incoming frames while generating; a synthesized interaction corpus supervises when to speak, when to stay silent, and when to revise; and a staged curriculum concentrates all real-time-specific training in one light final stage over a strong offline foundation. Offline, MOSS-VL-Instruct is competitive at comparable scale and leads temporal-reasoning video sets. Across four streaming benchmarks, MOSS-VL-Realtime posts the best average on three (second on the fourth) among open-source streaming models, sweeping the three subsets that squarely test proactive behavior -- 66.0 vs. 37.5 for the best baseline on OmniMMI Proactive Alerting. With 11.3B parameters but visual tokens outside the decoded sequence, MOSS-VL widens its time-to-first-token advantage over same-backbone Qwen3-VL-8B from 2.8x to 5.1x as visual context grows. We release all five checkpoints, the training curriculum, and the real-time inference code at https://github.com/OpenMOSS/MOSS-VL.
Problem

Research questions and friction points this paper is trying to address.

Vision-Language Model
Real-time Interaction
Streaming
Proactive Behavior
Latency
Innovation

Methods, ideas, or system contributions that make the work stand out.

Gated Cross-Attention
Real-time Interaction
Staged Curriculum Learning
Synthesized Interaction Corpus
Low Latency Inference
💼 Related Jobs
No related jobs found.
P
Pengyu Wang
Chenkun Tan
Chenkun Tan
Fudan University
Shaojun Zhou
Shaojun Zhou
Fudan University
Q
Qirui Zhou
Y
Yanxin Chen
X
Xingyang He
H
Huazheng Zeng
J
Jijun Cheng
Chenghao Wang
Chenghao Wang
Northeastern University
Robotics
X
Xiaomeng Qian
P
Pengfei Wang
Z
Zhan Huang
S
Shanqing Gao
W
Wei Huang
L
Longjun Cao
W
Wu Ran
J
Jie Liu
C
Changtai Zhu
H
Hongkai Wang
Y
Yixian Tian
C
Chenghao Liu
Z
Zhen Ye
Xinghao Wang
Xinghao Wang
Fudan University
Natural Language ProcessingLarge Language Models
B
Botian Jiang
G
Guoguo Feng