EchoWM: Open and Enterable Omnimodal World Models

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出EchoWM模型,通过共享的6-DoF轨迹和数据校准方法,解决连续导航下生成高质量视频、声音及语音的问题。
📝 Abstract
We present EchoWM, an omnimodal world model for enterable generative media that responds to continuous navigation while jointly generating 720p video, environmental sound, music and speech. We organize interaction around camera intent: in first-person scenes, it specifies observer motion, while in third-person scenes, camera--character dynamics are learned from data without view-specific controllers. Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data. To jointly learn audio-visual generation and trajectory control, we construct a complementary data engine and adopt progressive training followed by autoregressive post-training for long-horizon generation. Extensive evaluations show that \model achieves strong trajectory following and high visual quality on public world-model benchmarks, supporting both first- and third-person interaction across varied subjects, and maintaining synchronized environmental sound and speech over long-horizon generation.
Problem

Research questions and friction points this paper is trying to address.

omnimodal world model
enterable generative media
continuous navigation
audio-visual generation
Innovation

Methods, ideas, or system contributions that make the work stand out.

omnimodal world model
enterable generative media
camera intent
trajectory control
progressive training
🔎 Similar Papers
No similar papers found.