Institution profile

Guangzhou Quwan Network Technology Co. Ltd

Industry researchasia · cn
Research library7linked papers
Opportunities0open roles
Selected work

Representative Papers

High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer

Dec 26, 2025

Existing methods struggle to simultaneously preserve fine-grained facial and hand details while maintaining spatiotemporal consistency in long-duration (>3-second) human image animation. To address this, we propose a high-fidelity long-video generation framework based on the Diffusion Transformer (DiT). Our approach introduces three key innovations: (1) a novel hybrid implicit guidance signal combined with a sharpness-aware guidance factor; (2) a temporal-aware positional offset adaptation module enabling arbitrary-length video synthesis; and (3) skeleton-aligned modeling coupled with identity-agnostic data augmentation. These components collectively enhance fine-grained structural modeling and inter-frame coherence. Quantitative and qualitative evaluations demonstrate state-of-the-art performance in critical metrics—including facial expression fidelity, hand motion dynamics, and temporal smoothness—while achieving superior visual quality and strong spatiotemporal consistency.

0 citationsRead paper

Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language Models

Nov 12, 2025

Existing LLM evaluation for emotional support relies predominantly on static, short-turn dialogues, failing to capture the dynamic evolution and longitudinal nature of human emotions. To address this, we propose the first evaluation framework explicitly designed for emotional dynamic trajectories. Our method introduces a first-order Markov emotional trajectory model, integrating psychologically grounded mechanisms—including contextual selection and cognitive reappraisal—to generate a large-scale benchmark comprising 328 emotion scenarios and 1,152 distractor events. We design a trajectory-level evaluation paradigm featuring constraints on emotion regulation strategies and causal adjustment for emotional state tracking. Furthermore, we introduce three novel metrics: BEL (Emotional Baseline Shift), ETV (Emotional Trajectory Variance), and ECP (Empathic Consistency Probability). Extensive evaluation across diverse LLMs demonstrates that our framework effectively discriminates long-term emotional support capabilities, yielding interpretable and actionable assessments of empathic interaction.

0 citationsRead paper

Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback

Oct 13, 2025

Existing audio-driven video generation methods suffer from critical limitations in lip-sync accuracy, long-term temporal coherence, and multi-character coordinated animation. To address these, we propose Mask-CFG—a training-free framework integrating positional offset inference, LoRA-based efficient fine-tuning, reward-guided optimization, and block-wise diffusion Transformer (DiT) inference—enabling high-fidelity, long-duration, natural dialogue video synthesis for arbitrary numbers of characters without architectural modification or domain-specific data. Our core innovation lies in decoupling character control from speech-driven motion via masked classifier-free guidance and dynamic reward calibration. This significantly improves lip-sync accuracy (+12.7% LSE), temporal consistency (+38.4% TCV), and multi-character motion naturalness. On multi-character benchmarks, Mask-CFG surpasses state-of-the-art methods while maintaining high fidelity, low inference cost, and strong generalization.

0 citationsRead paper

HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS

Sep 23, 2025

Existing instructable TTS models suffer from a modality gap between coarse-grained text instructions and fine-grained speech tokens, hindering precise prosodic and phonetic control. To address this, we propose HD-PPT—a Hierarchical Dual-Preference Prompting and Tokenization framework. First, we design a novel hierarchical speech codec that disentangles content-related and instruction-related speech tokens. Second, we introduce a dual-preference token extraction mechanism jointly supervised by ASR and CLAP to align textual instructions with hierarchical speech representations. Third, we establish a layered decoding process enabling controllable generation across semantic, prosodic, and phonemic levels. By integrating large language models, speech codec architectures, ASR, CLAP, and hierarchical modeling, HD-PPT achieves state-of-the-art performance in both instruction adherence and speech naturalness, significantly improving the accuracy and expressiveness of controllable TTS.

0 citationsRead paper

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Aug 06, 2025

Traditional ASR and TTS systems typically neglect paralinguistic vocalizations—such as laughter, breathing, and filled pauses (“uh”, “oh”)—despite their critical role in emotional expression and conversational interaction. To address this, we propose the first end-to-end paralinguistic-aware Mandarin speech modeling framework: (1) We formulate paralinguistic information as learnable, decodable head tokens, enabling unified paralinguistic recognition and controllable synthesis; (2) We construct the first large-scale word-level annotated Chinese paralinguistic speech dataset—comprising 48k manually and 174k automatically labeled utterances (573 hours total); (3) Leveraging this dataset, we train a paralinguistic-aware ASR system and fine-tune a zero-shot TTS model to generate context-aware paralinguistic vocalizations. Experiments demonstrate substantial improvements in speech naturalness, expressiveness, and controllability over baseline systems.

0 citationsRead paper
Recent publications

Latest Papers

High-Fidelity and Long-Duration Human Image Animation with Diffusion Transformer

Dec 26, 2025

Existing methods struggle to simultaneously preserve fine-grained facial and hand details while maintaining spatiotemporal consistency in long-duration (>3-second) human image animation. To address this, we propose a high-fidelity long-video generation framework based on the Diffusion Transformer (DiT). Our approach introduces three key innovations: (1) a novel hybrid implicit guidance signal combined with a sharpness-aware guidance factor; (2) a temporal-aware positional offset adaptation module enabling arbitrary-length video synthesis; and (3) skeleton-aligned modeling coupled with identity-agnostic data augmentation. These components collectively enhance fine-grained structural modeling and inter-frame coherence. Quantitative and qualitative evaluations demonstrate state-of-the-art performance in critical metrics—including facial expression fidelity, hand motion dynamics, and temporal smoothness—while achieving superior visual quality and strong spatiotemporal consistency.

0 citationsRead paper

Detecting Emotional Dynamic Trajectories: An Evaluation Framework for Emotional Support in Language Models

Nov 12, 2025

Existing LLM evaluation for emotional support relies predominantly on static, short-turn dialogues, failing to capture the dynamic evolution and longitudinal nature of human emotions. To address this, we propose the first evaluation framework explicitly designed for emotional dynamic trajectories. Our method introduces a first-order Markov emotional trajectory model, integrating psychologically grounded mechanisms—including contextual selection and cognitive reappraisal—to generate a large-scale benchmark comprising 328 emotion scenarios and 1,152 distractor events. We design a trajectory-level evaluation paradigm featuring constraints on emotion regulation strategies and causal adjustment for emotional state tracking. Furthermore, we introduce three novel metrics: BEL (Emotional Baseline Shift), ETV (Emotional Trajectory Variance), and ECP (Empathic Consistency Probability). Extensive evaluation across diverse LLMs demonstrates that our framework effectively discriminates long-term emotional support capabilities, yielding interpretable and actionable assessments of empathic interaction.

0 citationsRead paper

Playmate2: Training-Free Multi-Character Audio-Driven Animation via Diffusion Transformer with Reward Feedback

Oct 13, 2025

Existing audio-driven video generation methods suffer from critical limitations in lip-sync accuracy, long-term temporal coherence, and multi-character coordinated animation. To address these, we propose Mask-CFG—a training-free framework integrating positional offset inference, LoRA-based efficient fine-tuning, reward-guided optimization, and block-wise diffusion Transformer (DiT) inference—enabling high-fidelity, long-duration, natural dialogue video synthesis for arbitrary numbers of characters without architectural modification or domain-specific data. Our core innovation lies in decoupling character control from speech-driven motion via masked classifier-free guidance and dynamic reward calibration. This significantly improves lip-sync accuracy (+12.7% LSE), temporal consistency (+38.4% TCV), and multi-character motion naturalness. On multi-character benchmarks, Mask-CFG surpasses state-of-the-art methods while maintaining high fidelity, low inference cost, and strong generalization.

0 citationsRead paper

HD-PPT: Hierarchical Decoding of Content- and Prompt-Preference Tokens for Instruction-based TTS

Sep 23, 2025

Existing instructable TTS models suffer from a modality gap between coarse-grained text instructions and fine-grained speech tokens, hindering precise prosodic and phonetic control. To address this, we propose HD-PPT—a Hierarchical Dual-Preference Prompting and Tokenization framework. First, we design a novel hierarchical speech codec that disentangles content-related and instruction-related speech tokens. Second, we introduce a dual-preference token extraction mechanism jointly supervised by ASR and CLAP to align textual instructions with hierarchical speech representations. Third, we establish a layered decoding process enabling controllable generation across semantic, prosodic, and phonemic levels. By integrating large language models, speech codec architectures, ASR, CLAP, and hierarchical modeling, HD-PPT achieves state-of-the-art performance in both instruction adherence and speech naturalness, significantly improving the accuracy and expressiveness of controllable TTS.

0 citationsRead paper

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations

Aug 06, 2025

Traditional ASR and TTS systems typically neglect paralinguistic vocalizations—such as laughter, breathing, and filled pauses (“uh”, “oh”)—despite their critical role in emotional expression and conversational interaction. To address this, we propose the first end-to-end paralinguistic-aware Mandarin speech modeling framework: (1) We formulate paralinguistic information as learnable, decodable head tokens, enabling unified paralinguistic recognition and controllable synthesis; (2) We construct the first large-scale word-level annotated Chinese paralinguistic speech dataset—comprising 48k manually and 174k automatically labeled utterances (573 hours total); (3) Leveraging this dataset, we train a paralinguistic-aware ASR system and fine-tune a zero-shot TTS model to generate context-aware paralinguistic vocalizations. Experiments demonstrate substantial improvements in speech naturalness, expressiveness, and controllability over baseline systems.

0 citationsRead paper