Institution profile

Movian AI

Industry research
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models

Jul 18, 2025

This work addresses the single-image content-style decomposition (CSD) problem, aiming for high-fidelity content extraction and controllable style transfer. We propose a scale-aware disentanglement framework grounded in visual autoregressive modeling: (i) a scale-aware alternating optimization strategy to strengthen content-style separation; (ii) an SVD-driven style rectification module to suppress content leakage into style representations; and (iii) an enhanced key-value memory mechanism to improve identity consistency across stylized outputs. To further boost generalization, we introduce scale-aligned training and targeted data augmentation. Additionally, we release CSD-100—the first dedicated benchmark for CSD evaluation. Extensive experiments demonstrate that our method achieves significant improvements over state-of-the-art approaches in both content fidelity and style consistency, enabling more flexible and precise visual synthesis with enhanced controllability.

0 citationsRead paper

Zero-Shot Text-to-Speech for Vietnamese

Jun 02, 2025

This study addresses the limited performance of zero-shot text-to-speech (TTS) for Vietnamese—a low-resource language—by constructing and open-sourcing PhoAudiobook, the first large-scale, high-quality Vietnamese speech dataset (941 hours), specifically designed for zero-shot TTS evaluation and training. Leveraging PhoAudiobook, we systematically benchmark three state-of-the-art cross-lingual TTS models—VALL-E, VoiceCraft, and XTTS-V2—on Vietnamese synthesis. Our experiments reveal, for the first time, that VALL-E and VoiceCraft exhibit strong cross-lingual robustness in short-utterance synthesis; moreover, PhoAudiobook consistently improves all models across critical metrics—including naturalness, speaker similarity, and intelligibility. This work fills a critical gap by establishing the first high-fidelity, zero-shot TTS benchmark for Vietnamese and provides a reproducible, open-source infrastructure and empirical foundation for zero-shot TTS research in low-resource languages.

0 citationsRead paper
Recent publications

Latest Papers

CSD-VAR: Content-Style Decomposition in Visual Autoregressive Models

Jul 18, 2025

This work addresses the single-image content-style decomposition (CSD) problem, aiming for high-fidelity content extraction and controllable style transfer. We propose a scale-aware disentanglement framework grounded in visual autoregressive modeling: (i) a scale-aware alternating optimization strategy to strengthen content-style separation; (ii) an SVD-driven style rectification module to suppress content leakage into style representations; and (iii) an enhanced key-value memory mechanism to improve identity consistency across stylized outputs. To further boost generalization, we introduce scale-aligned training and targeted data augmentation. Additionally, we release CSD-100—the first dedicated benchmark for CSD evaluation. Extensive experiments demonstrate that our method achieves significant improvements over state-of-the-art approaches in both content fidelity and style consistency, enabling more flexible and precise visual synthesis with enhanced controllability.

0 citationsRead paper

Zero-Shot Text-to-Speech for Vietnamese

Jun 02, 2025

This study addresses the limited performance of zero-shot text-to-speech (TTS) for Vietnamese—a low-resource language—by constructing and open-sourcing PhoAudiobook, the first large-scale, high-quality Vietnamese speech dataset (941 hours), specifically designed for zero-shot TTS evaluation and training. Leveraging PhoAudiobook, we systematically benchmark three state-of-the-art cross-lingual TTS models—VALL-E, VoiceCraft, and XTTS-V2—on Vietnamese synthesis. Our experiments reveal, for the first time, that VALL-E and VoiceCraft exhibit strong cross-lingual robustness in short-utterance synthesis; moreover, PhoAudiobook consistently improves all models across critical metrics—including naturalness, speaker similarity, and intelligibility. This work fills a critical gap by establishing the first high-fidelity, zero-shot TTS benchmark for Vietnamese and provides a reproducible, open-source infrastructure and empirical foundation for zero-shot TTS research in low-resource languages.

0 citationsRead paper