Realtime-Venus: A full-duplex interaction system with asynchronous delegation

📅 2026-09-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为了解决数字和物理环境中自然交互的连续感知和及时响应问题,通过两个分别训练的9B模型Realtime-Venus-Omni和Realtime-Venus-Audio来支持视听与语音互动。
📝 Abstract
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events. A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue. Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows. Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.
Problem

Research questions and friction points this paper is trying to address.

natural interaction
continuous perception
timely responses
spoken dialogue
video interaction
Innovation

Methods, ideas, or system contributions that make the work stand out.

full-duplex interaction
asynchronous delegation
audio-visual interaction
spoken interaction
continuous perception
Ruixiang Zhao
Ruixiang Zhao
Renmin University of China
H
Hualei Wang
Ant Group
R
Renhe Sun
Ant Group
E
Enzhi Zhou
Ant Group
Jincenzi Wu
Jincenzi Wu
The Chinese University of Hong Kong
Natural Language ProcessingAI AnthropomorphismComputational Psychology
X
Xujie Song
Ant Group
K
Kexin Shi
Ant Group
Z
Zihang Liu
Ant Group
Pengcheng Zhu
Pengcheng Zhu
Fuxi AI Lab, NetEase Inc.
speech synthesissinging voice synthesistalking avatarvoice conversion
J
Jiayi Zhou
Ant Group
Baoyue Zhang
Baoyue Zhang
Ant Group
C
Changhao Zhang
Ant Group
Z
Zitong Wang
Ant Group
J
Jinhong Wang
Ant Group
T
Tong Niu
Ant Group
J
Jingjing Liu
Ant Group
J
Junan Lin
Ant Group
H
Haolin He
Ant Group
H
Hengshuo Chu
Ant Group
Y
Yuhui Chen
Ant Group
J
Jian Liu
Ant Group
Y
Yuge Huang
Ant Group
J
Junliang Xing
Ant Group
Yuntao Wang
Yuntao Wang
Tsinghua University
Human-Computer InteractionUbiquitous ComputingPhysio-Behavioral Computing
W
Weiqiang Wang
Ant Group