Omni Interaction Agent Technical Report

📅 2026-09-08
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出Gander模型,通过脑-小脑协作框架和流式思考-说话架构解决全感知实时互动问题,支持多模态输入和连续自然对话。
📝 Abstract
In this work, we present Gander, an end-to-end model that unifies omni perception, realtime interaction, and agentic capabilities within a single framework. In contrast to turn-based conventional paradigms, Gander continuously receives streaming inputs across multiple modalities, including video, speech, and text, enabling natural full-duplex interaction in both everyday conversations and complex workflow-oriented agent scenarios. Users can interrupt the model at any time, while the model can also proactively provide intermediate feedback or ask follow up questions. To natively support these capabilities, Gander adopts two key architectural designs: 1) It employs a Cerebellum-Brain collaborative framework, in which the Cerebellum is responsible for realtime interaction and omni conversational capabilities, while the Brain handles complex reasoning and higher-level agentic tasks. The two components interact continuously through tool calling and the agent orchestration runtime. 2) The Cerebellum is built upon a streaming Thinker-Talker architecture, user inputs and model outputs are further flattened into an ordered token stream at the chunk level, providing a unified representation for low latency, continuous interaction. We conduct comprehensive evaluations of Gander across four dimensions: conversational ability, omni understanding, interactive capability, and agentic intelligence. Internal human evaluations demonstrate that Gander maintains the natural and expressive spoken dialogue capabilities of SOTA open source models while achieving competitive performance in omni interaction. Gander also demonstrates robustness in challenging real-world scenarios, including background noise interference, multi-party interactions, and backchannel communication. We release Gander together with its models, code, and data to facilitate further research and development in the community.
Problem

Research questions and friction points this paper is trying to address.

omni perception
realtime interaction
full-duplex interaction
multi-modal inputs
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end model
omni perception
realtime interaction
Cerebellum-Brain collaborative framework
streaming Thinker-Talker architecture
💼 Related Jobs
No related jobs found.
O
Orantqing
Hunyuan Speech Team, Tencent
S
Shengpeng Ji
Hunyuan Speech Team, Tencent; Zhejiang University
J
Junlong Tong
Hunyuan Speech Team, Tencent; Shanghai Jiao Tong University
Jialong Zuo
Jialong Zuo
Zhejiang University
Speech SynthesisVoice Conversion
Dongjie Fu
Dongjie Fu
LREIS, Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences
Remote Sensing
Di Cao
Di Cao
University of Electronic Science and Technology of China
power systemmachine learning
Y
Yangzhuo Li
Hunyuan Speech Team, Tencent
Shangda Wu
Shangda Wu
Tencent
Symbolic Music GenerationMusic Information RetrievalMultimodal Learning
F
Franz
Hunyuan Speech Team, Tencent
E
Evan
Hunyuan Speech Team, Tencent
T
Theron Veyra
Hunyuan Speech Team, Tencent
Changhao Pan
Changhao Pan
Zhejiang University
Multi-Modal Genarative AISinging Voice Synthesis
J
Jingyu Lu
Zhejiang University
Dongchao Yang
Dongchao Yang
Chinese University of Hong Kong
TTSTTAAudio CodecMulti-modal Audio Fundation Models
Zhifei Xie
Zhifei Xie
Tsinghua University
Artificial IntelligenceLarge Multimodal ModelGPT4o
Y
Yang Tan
Hunyuan Speech Team, Tencent
Xiaoyu Shen
Xiaoyu Shen
Eastern Institute of Technology, Ningbo
language modelmulti-modal learningreasoning
X
Xiaoda Yang
Zhejiang University
W
Wenfu Wang
Hunyuan Speech Team, Tencent
T
Teddy Sun
Hunyuan Speech Team, Tencent
S
Steve Yves
Hunyuan Speech Team, Tencent
Zhou Zhao
Zhou Zhao
Zhejiang University
Machine LearningData MiningMultimedia Computing