Institution profile

Hedra

Industry researchnorthamerica · us
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching

Jul 21, 2025

To address inaccurate lip synchronization and long-term pose drift in real-time audio-driven talking-head video generation, this paper proposes a real-time-optimized tailored flow matching framework. Methodologically, it integrates audio-feature conditioning, explicit pose modeling, and efficient inference optimization to enable end-to-end low-latency sequence generation. Its key contribution lies in a lightweight flow matching architecture that significantly improves visual naturalness and temporal stability while preserving high frame-level temporal coherence. Experiments on the HDTF dataset demonstrate a LipSync Confidence score of 8.50, an inference throughput of 141 FPS on a single A10 GPU, and an end-to-end latency of only 0.17 seconds—enabling high-fidelity virtual avatar deployment across diverse real-time scenarios.

0 citationsRead paper
Recent publications

Latest Papers

Livatar-1: Real-Time Talking Heads Generation with Tailored Flow Matching

Jul 21, 2025

To address inaccurate lip synchronization and long-term pose drift in real-time audio-driven talking-head video generation, this paper proposes a real-time-optimized tailored flow matching framework. Methodologically, it integrates audio-feature conditioning, explicit pose modeling, and efficient inference optimization to enable end-to-end low-latency sequence generation. Its key contribution lies in a lightweight flow matching architecture that significantly improves visual naturalness and temporal stability while preserving high frame-level temporal coherence. Experiments on the HDTF dataset demonstrate a LipSync Confidence score of 8.50, an inference throughput of 141 FPS on a single A10 GPU, and an end-to-end latency of only 0.17 seconds—enabling high-fidelity virtual avatar deployment across diverse real-time scenarios.

0 citationsRead paper