Institution profile

Shanghai Soulgate Techonolgy Co.tl

Industry researchasia · cn
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation

Mar 24, 2025

This work addresses two key challenges in real-time audio-driven portrait animation: difficulty in long-term temporal modeling and inconsistent multi-region motion. To this end, we propose the first autoregressive framework specifically designed for streaming speech. Methodologically: (1) we design an audio-stream-driven autoregressive facial motion token generation mechanism; (2) we introduce implicit keypoint modeling coupled with an Efficient Temporal Module (ETM) to explicitly capture subtle physical motions—such as neck muscle deformation and earring oscillation; and (3) we integrate Residual Vector Quantization (Residual VQ) with a lightweight Transformer architecture to enhance computational efficiency. Experiments demonstrate real-time generation at 25 FPS, with inference latency of only 0.92 seconds per second of video—22× faster than diffusion-based methods. Human evaluation confirms significant improvements in fine-motion fidelity and overall visual realism.

0 citationsRead paper
Recent publications

Latest Papers

Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation

Mar 24, 2025

This work addresses two key challenges in real-time audio-driven portrait animation: difficulty in long-term temporal modeling and inconsistent multi-region motion. To this end, we propose the first autoregressive framework specifically designed for streaming speech. Methodologically: (1) we design an audio-stream-driven autoregressive facial motion token generation mechanism; (2) we introduce implicit keypoint modeling coupled with an Efficient Temporal Module (ETM) to explicitly capture subtle physical motions—such as neck muscle deformation and earring oscillation; and (3) we integrate Residual Vector Quantization (Residual VQ) with a lightweight Transformer architecture to enhance computational efficiency. Experiments demonstrate real-time generation at 25 FPS, with inference latency of only 0.92 seconds per second of video—22× faster than diffusion-based methods. Human evaluation confirms significant improvements in fine-motion fidelity and overall visual realism.

0 citationsRead paper