Cross-Lingual F5-TTS 2: A Simplified Framework for Language-Agnostic Voice Cloning
本文提出Cross-Lingual F5-TTS 2,通过预训练模型和细调解决无文本跨语言声音克隆问题,避免了强制对齐的复杂性和误差。
本文提出Cross-Lingual F5-TTS 2,通过预训练模型和细调解决无文本跨语言声音克隆问题,避免了强制对齐的复杂性和误差。
Existing prompt inversion methods for text-to-image diffusion models struggle to simultaneously preserve image fidelity and semantic interpretability, often overlooking the critical role of latent noise in structural consistency. This work proposes Dualin, a two-stage joint inversion framework: the first stage integrates vision-language and large language models to generate faithful and human-readable hard prompts, while the second stage employs unconditional DDIM inversion to accurately recover latent noise, achieving dual alignment in both semantics and structure. Dualin is the first approach to unify prompt inversion with latent noise reconstruction, and we theoretically demonstrate that the recovered noise enables flexible editing without re-optimization, overcoming the limitations of prompt-only paradigms. Experiments show that Dualin achieves high-quality prompt inversion and state-of-the-art image fidelity across multiple datasets, laying a foundation for precise and controllable image editing.
Existing approaches treat controllable generation with diffusion models and dense prediction as disjoint tasks, overlooking the potential of jointly modeling their heterogeneous output distributions. This work proposes the Dense Unified Generation and Perception (DUGP) module and a unified dataset training strategy built upon the MMDiT architecture, enabling non-RGB dense outputs to be modeled simply by duplicating the image branch—thereby achieving mutual enhancement between generative and perceptual tasks. Adhering to Occam’s razor, the method requires no task-specific designs and instead learns the joint distribution of image-geometry pairs through multi-task co-training. The resulting unified model surpasses prior unified approaches and matches the performance of specialized models: generative priors enrich perceptual detail, while perception signals improve structural alignment in generation.
This work addresses the limitations of traditional SLAM, which lacks predictive planning capabilities, and action-conditioned world models, which often disregard metric consistency and suffer from open-loop drift. The paper presents the first unified formulation of SLAM and action-conditioned world modeling as a joint estimation problem, introducing a coupled factor graph framework. This framework enables bidirectional optimization between localization and actionable world representation through pose-latent coupling factors and supports latent loop closure. By integrating multi-source probabilistic constraints—including observations, action predictions, odometry, latent landmarks, and loop closures—the approach significantly reduces latent prediction RMSE and end-point drift on real-world PushT and WildGS datasets, while generating globally consistent maps that preserve both metric accuracy and actionability.
Traditional large language models exhibit limited generalization under corpus heterogeneity and fine-grained conditional shifts, while fine-tuning often incurs catastrophic forgetting, and existing meta-learning approaches struggle to scale to large models. This work proposes a novel meta-control paradigm that introduces a learnable β parameter within the SwiGLU module as a meta-signal, enabling an adaptive meta-gating mechanism to modulate the nonlinearity of the feedforward network. A hypernetwork is further designed to dynamically generate β parameters conditioned on arbitrary textual attributes—such as task, domain, persona, or style—thereby unifying multidimensional textual conditions into a single, generalizable control framework. The method achieves state-of-the-art performance across both seen and unseen conditions, significantly outperforming standard fine-tuning and existing meta-learning baselines.
本文提出Cross-Lingual F5-TTS 2,通过预训练模型和细调解决无文本跨语言声音克隆问题,避免了强制对齐的复杂性和误差。
Existing prompt inversion methods for text-to-image diffusion models struggle to simultaneously preserve image fidelity and semantic interpretability, often overlooking the critical role of latent noise in structural consistency. This work proposes Dualin, a two-stage joint inversion framework: the first stage integrates vision-language and large language models to generate faithful and human-readable hard prompts, while the second stage employs unconditional DDIM inversion to accurately recover latent noise, achieving dual alignment in both semantics and structure. Dualin is the first approach to unify prompt inversion with latent noise reconstruction, and we theoretically demonstrate that the recovered noise enables flexible editing without re-optimization, overcoming the limitations of prompt-only paradigms. Experiments show that Dualin achieves high-quality prompt inversion and state-of-the-art image fidelity across multiple datasets, laying a foundation for precise and controllable image editing.
Existing approaches treat controllable generation with diffusion models and dense prediction as disjoint tasks, overlooking the potential of jointly modeling their heterogeneous output distributions. This work proposes the Dense Unified Generation and Perception (DUGP) module and a unified dataset training strategy built upon the MMDiT architecture, enabling non-RGB dense outputs to be modeled simply by duplicating the image branch—thereby achieving mutual enhancement between generative and perceptual tasks. Adhering to Occam’s razor, the method requires no task-specific designs and instead learns the joint distribution of image-geometry pairs through multi-task co-training. The resulting unified model surpasses prior unified approaches and matches the performance of specialized models: generative priors enrich perceptual detail, while perception signals improve structural alignment in generation.
This work addresses the limitations of traditional SLAM, which lacks predictive planning capabilities, and action-conditioned world models, which often disregard metric consistency and suffer from open-loop drift. The paper presents the first unified formulation of SLAM and action-conditioned world modeling as a joint estimation problem, introducing a coupled factor graph framework. This framework enables bidirectional optimization between localization and actionable world representation through pose-latent coupling factors and supports latent loop closure. By integrating multi-source probabilistic constraints—including observations, action predictions, odometry, latent landmarks, and loop closures—the approach significantly reduces latent prediction RMSE and end-point drift on real-world PushT and WildGS datasets, while generating globally consistent maps that preserve both metric accuracy and actionability.
Traditional large language models exhibit limited generalization under corpus heterogeneity and fine-grained conditional shifts, while fine-tuning often incurs catastrophic forgetting, and existing meta-learning approaches struggle to scale to large models. This work proposes a novel meta-control paradigm that introduces a learnable β parameter within the SwiGLU module as a meta-signal, enabling an adaptive meta-gating mechanism to modulate the nonlinearity of the feedforward network. A hypernetwork is further designed to dynamically generate β parameters conditioned on arbitrary textual attributes—such as task, domain, persona, or style—thereby unifying multidimensional textual conditions into a single, generalizable control framework. The method achieves state-of-the-art performance across both seen and unseen conditions, significantly outperforming standard fine-tuning and existing meta-learning baselines.