Institution profile

Soul AIGC Team

Industry researchasia · cn
Research library4linked papers
Opportunities0open roles
Selected work

Representative Papers

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

Aug 07, 2026

This work addresses the limited content fidelity in existing continuous latent autoregressive speech generation methods, which lack explicit linguistic structural supervision. To remedy this, we propose SemBridge, a novel framework that introduces discrete semantic tokens—used only during training—to directly supervise the states of the autoregressive language model. A semantic-aligned acoustic variational autoencoder (VAE) is employed to construct a structured continuous target space. Notably, the inference pipeline remains fully continuous without any architectural modifications. The proposed approach substantially reduces word and character error rates while preserving high speaker similarity and perceptual quality. Furthermore, it enables zero-shot text-to-speech synthesis and score-conditioned singing voice generation.

0 citationsRead paper

SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads

Feb 07, 2026

Existing audio-driven talking-head generation methods struggle to simultaneously achieve high fidelity, temporal stability, and real-time streaming capability. This work proposes SoulX-FlashHead, a unified 1.3B-parameter framework enabling infinite-duration, high-fidelity, and low-latency talking-head video synthesis. Key innovations include stream-aware spatiotemporal pretraining, a temporal audio context caching mechanism to enhance the robustness of audio features, and a novel oracle-guided bidirectional distillation strategy that effectively mitigates error accumulation and identity drift in autoregressive generation. The model achieves state-of-the-art performance on both HDTF and VFHQ benchmarks. Its Lite variant attains 96 FPS inference on a single RTX 4090 GPU, striking an optimal balance between visual quality and ultra-low-latency interactivity.

0 citationsRead paper

Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression

Jun 11, 2025

This work addresses two key limitations in image generation: the low inference efficiency of autoregressive (AR) models and the weak semantic modeling capability of diffusion models. To this end, we propose TransDiff—the first unified framework that synergistically integrates autoregressive Transformers with diffusion modeling. Its core contributions are: (1) a Multi-Reference Autoregressive (MRAR) paradigm that conditions sequence modeling on multiple previously generated images, thereby enhancing diversity and fidelity; and (2) joint optimization of semantic feature encoding and diffusion-based distribution estimation for high-fidelity synthesis. On ImageNet 256×256, TransDiff achieves state-of-the-art performance with an FID of 1.42 and an Inception Score (IS) of 293.4. Moreover, it attains a 2× speedup over AR models and a 112× acceleration over standard diffusion models—demonstrating unprecedented balance among generation quality, semantic coherence, and inference efficiency.

0 citationsRead paper

Global Position Aware Group Choreography using Large Language Model

Mar 12, 2025

Prior research on group dance generation is scarce, and existing single-dancer methods do not scale effectively to multi-agent coordination. Method: We propose the first large language model (LLM)-based, music-driven group dance generation framework. It formalizes group choreography as an audio-to-multi-dancer motion sequence translation task, introducing position-aware tokenization and inter-dancer kinematic consistency constraints to jointly model individual expressivity and collective coordination while preserving audio synchronization. Our approach integrates continuous feature discretization via a learned tokenizer, LLM fine-tuning, multimodal sequence modeling, and physics-informed dance constraints. Contribution/Results: The method achieves state-of-the-art performance across multiple benchmarks, supports real-time co-generation for four or more dancers, and significantly improves musical alignment, spatiotemporal coherence, and visual diversity of generated group dances.

0 citationsRead paper
Recent publications

Latest Papers

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation

Aug 07, 2026

This work addresses the limited content fidelity in existing continuous latent autoregressive speech generation methods, which lack explicit linguistic structural supervision. To remedy this, we propose SemBridge, a novel framework that introduces discrete semantic tokens—used only during training—to directly supervise the states of the autoregressive language model. A semantic-aligned acoustic variational autoencoder (VAE) is employed to construct a structured continuous target space. Notably, the inference pipeline remains fully continuous without any architectural modifications. The proposed approach substantially reduces word and character error rates while preserving high speaker similarity and perceptual quality. Furthermore, it enables zero-shot text-to-speech synthesis and score-conditioned singing voice generation.

0 citationsRead paper

SoulX-FlashHead: Oracle-guided Generation of Infinite Real-time Streaming Talking Heads

Feb 07, 2026

Existing audio-driven talking-head generation methods struggle to simultaneously achieve high fidelity, temporal stability, and real-time streaming capability. This work proposes SoulX-FlashHead, a unified 1.3B-parameter framework enabling infinite-duration, high-fidelity, and low-latency talking-head video synthesis. Key innovations include stream-aware spatiotemporal pretraining, a temporal audio context caching mechanism to enhance the robustness of audio features, and a novel oracle-guided bidirectional distillation strategy that effectively mitigates error accumulation and identity drift in autoregressive generation. The model achieves state-of-the-art performance on both HDTF and VFHQ benchmarks. Its Lite variant attains 96 FPS inference on a single RTX 4090 GPU, striking an optimal balance between visual quality and ultra-low-latency interactivity.

0 citationsRead paper

Marrying Autoregressive Transformer and Diffusion with Multi-Reference Autoregression

Jun 11, 2025

This work addresses two key limitations in image generation: the low inference efficiency of autoregressive (AR) models and the weak semantic modeling capability of diffusion models. To this end, we propose TransDiff—the first unified framework that synergistically integrates autoregressive Transformers with diffusion modeling. Its core contributions are: (1) a Multi-Reference Autoregressive (MRAR) paradigm that conditions sequence modeling on multiple previously generated images, thereby enhancing diversity and fidelity; and (2) joint optimization of semantic feature encoding and diffusion-based distribution estimation for high-fidelity synthesis. On ImageNet 256×256, TransDiff achieves state-of-the-art performance with an FID of 1.42 and an Inception Score (IS) of 293.4. Moreover, it attains a 2× speedup over AR models and a 112× acceleration over standard diffusion models—demonstrating unprecedented balance among generation quality, semantic coherence, and inference efficiency.

0 citationsRead paper

Global Position Aware Group Choreography using Large Language Model

Mar 12, 2025

Prior research on group dance generation is scarce, and existing single-dancer methods do not scale effectively to multi-agent coordination. Method: We propose the first large language model (LLM)-based, music-driven group dance generation framework. It formalizes group choreography as an audio-to-multi-dancer motion sequence translation task, introducing position-aware tokenization and inter-dancer kinematic consistency constraints to jointly model individual expressivity and collective coordination while preserving audio synchronization. Our approach integrates continuous feature discretization via a learned tokenizer, LLM fine-tuning, multimodal sequence modeling, and physics-informed dance constraints. Contribution/Results: The method achieves state-of-the-art performance across multiple benchmarks, supports real-time co-generation for four or more dancers, and significantly improves musical alignment, spatiotemporal coherence, and visual diversity of generated group dances.

0 citationsRead paper