Institution profile

Mogo Auto Intelligence and Telemetics Information Technology Co. Ltd.

Industry researchasia · cn
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

MOGO: Residual Quantized Hierarchical Causal Transformer for High-Quality and Real-Time 3D Human Motion Generation

Jun 06, 2025

This work addresses the challenge of simultaneously achieving high fidelity, low latency, streaming generation, and scalability in text-driven 3D human motion synthesis. We propose a joint framework comprising MoSA-VQ (Motion Scale-Adaptive residual Vector Quantization) and RQHC-Transformer (Residual Quantized Hierarchical Causal Transformer). MoSA-VQ enables compact, multi-granularity motion representation via scale-adaptive residual quantization, while RQHC-Transformer supports single-step, multi-layer motion token generation and streaming autoregressive decoding. A novel text-motion cross-modal alignment mechanism is introduced to enhance semantic consistency. Evaluated on HumanML3D, KIT-ML, and CMP benchmarks, our method achieves state-of-the-art generation quality—measured by diversity, realism, and text-motion alignment—while significantly reducing inference latency. Notably, it is the first approach to enable real-time, high-fidelity, zero-shot streaming motion generation within a single forward pass.

0 citationsRead paper
Recent publications

Latest Papers

MOGO: Residual Quantized Hierarchical Causal Transformer for High-Quality and Real-Time 3D Human Motion Generation

Jun 06, 2025

This work addresses the challenge of simultaneously achieving high fidelity, low latency, streaming generation, and scalability in text-driven 3D human motion synthesis. We propose a joint framework comprising MoSA-VQ (Motion Scale-Adaptive residual Vector Quantization) and RQHC-Transformer (Residual Quantized Hierarchical Causal Transformer). MoSA-VQ enables compact, multi-granularity motion representation via scale-adaptive residual quantization, while RQHC-Transformer supports single-step, multi-layer motion token generation and streaming autoregressive decoding. A novel text-motion cross-modal alignment mechanism is introduced to enhance semantic consistency. Evaluated on HumanML3D, KIT-ML, and CMP benchmarks, our method achieves state-of-the-art generation quality—measured by diversity, realism, and text-motion alignment—while significantly reducing inference latency. Notably, it is the first approach to enable real-time, high-fidelity, zero-shot streaming motion generation within a single forward pass.

0 citationsRead paper