EraseSAE: Surgical Concept Erasure in Text-to-Video Diffusion Models via Sparse Autoencoders
为解决文本到视频扩散模型中的安全和版权问题,本文提出EraseSAE框架,利用稀疏自动编码器实现精准的概念移除。
为解决文本到视频扩散模型中的安全和版权问题,本文提出EraseSAE框架,利用稀疏自动编码器实现精准的概念移除。
本文提出RASA框架,通过解耦空间映射和运动控制解决跨身份角色动画中的位置、比例对齐及关节细化问题,显著提升动画的真实性和视觉质量。
Existing virtual try-on methods struggle to achieve fine-grained control over garment layout, often yielding results with limited diversity. This work proposes MOFA-VTON, the first approach to introduce a user sketch-driven dual-region masking mechanism combined with a cross-attention-based layout refinement module, enabling independent and precise spatial manipulation of upper- and lower-body garments. By integrating sketch-to-mask conversion with a region-aware generative network, MOFA-VTON overcomes the constraints of fixed-layout paradigms, facilitating interactive and high-fidelity virtual try-on. Extensive experiments on the VITON-HD and DressCode datasets demonstrate that MOFA-VTON significantly outperforms state-of-the-art methods, achieving notable improvements in outfit diversity, photorealism, and fashion expressiveness.
This work addresses key limitations of existing visual autoregressive (VAR) models in subject-driven image generation, particularly the inconsistency between multi-scale conditional training and inference as well as insufficient semantic alignment. To resolve these issues, the authors propose a pre-filled subject feature sequence mechanism that extracts reference subject features via a multi-scale visual tokenizer and fully injects them prior to autoregressive generation, thereby eliminating train-test discrepancies and simplifying dependency modeling. Furthermore, this study introduces reinforcement learning into the VAR framework for the first time, jointly optimizing semantic alignment and subject consistency. The proposed approach significantly enhances appearance fidelity and achieves superior generation quality compared to state-of-the-art diffusion models.
Text-guided image inpainting faces the challenge of simultaneously achieving prompt alignment and visual coherence. This paper proposes a fine-tuning-free, plug-and-play framework that directly optimizes latent variables of diffusion models during inference. Our method addresses this challenge through two key innovations: (1) a novel prior-guided optimization mechanism operating in the noise space, and (2) a composite guidance objective tailored for inpainting—integrating attention-region constraints with multi-step intermediate latent guidance. Crucially, the approach preserves the pre-trained model intact while jointly enhancing prompt fidelity and visual coherence. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across multiple benchmarks and quantitative metrics, exhibiting strong generalization capability and practical plug-and-play utility.
为解决文本到视频扩散模型中的安全和版权问题,本文提出EraseSAE框架,利用稀疏自动编码器实现精准的概念移除。
本文提出RASA框架,通过解耦空间映射和运动控制解决跨身份角色动画中的位置、比例对齐及关节细化问题,显著提升动画的真实性和视觉质量。
Existing virtual try-on methods struggle to achieve fine-grained control over garment layout, often yielding results with limited diversity. This work proposes MOFA-VTON, the first approach to introduce a user sketch-driven dual-region masking mechanism combined with a cross-attention-based layout refinement module, enabling independent and precise spatial manipulation of upper- and lower-body garments. By integrating sketch-to-mask conversion with a region-aware generative network, MOFA-VTON overcomes the constraints of fixed-layout paradigms, facilitating interactive and high-fidelity virtual try-on. Extensive experiments on the VITON-HD and DressCode datasets demonstrate that MOFA-VTON significantly outperforms state-of-the-art methods, achieving notable improvements in outfit diversity, photorealism, and fashion expressiveness.
This work addresses key limitations of existing visual autoregressive (VAR) models in subject-driven image generation, particularly the inconsistency between multi-scale conditional training and inference as well as insufficient semantic alignment. To resolve these issues, the authors propose a pre-filled subject feature sequence mechanism that extracts reference subject features via a multi-scale visual tokenizer and fully injects them prior to autoregressive generation, thereby eliminating train-test discrepancies and simplifying dependency modeling. Furthermore, this study introduces reinforcement learning into the VAR framework for the first time, jointly optimizing semantic alignment and subject consistency. The proposed approach significantly enhances appearance fidelity and achieves superior generation quality compared to state-of-the-art diffusion models.
Text-guided image inpainting faces the challenge of simultaneously achieving prompt alignment and visual coherence. This paper proposes a fine-tuning-free, plug-and-play framework that directly optimizes latent variables of diffusion models during inference. Our method addresses this challenge through two key innovations: (1) a novel prior-guided optimization mechanism operating in the noise space, and (2) a composite guidance objective tailored for inpainting—integrating attention-region constraints with multi-step intermediate latent guidance. Crucially, the approach preserves the pre-trained model intact while jointly enhancing prompt fidelity and visual coherence. Extensive experiments demonstrate that our method significantly outperforms state-of-the-art approaches across multiple benchmarks and quantitative metrics, exhibiting strong generalization capability and practical plug-and-play utility.