Ouroboros: Self-Referential Backdoor Attacks on Speech Enhancement via Clean Audio Triggers
本文提出Ouroboros框架,利用语音增强模型的理想纯净输出作为自然触发器,在不注入外部触发器的情况下实现后门攻击,适用于多种模型和数据集。
本文提出Ouroboros框架,利用语音增强模型的理想纯净输出作为自然触发器,在不注入外部触发器的情况下实现后门攻击,适用于多种模型和数据集。
本文提出GarmentWeaver,通过激活相关结构分支和采用预训练的视觉-语言模型来生成更准确且可执行的缝纫图样,解决了现有方法中结构与细节参数混淆的问题。
研究通过虚拟现实实验探讨了自动驾驶卡车与行人在多车道道路上的互动风险,并测试了三种干预措施,发现投影式外部人机界面最有效。
本文提出DAMOS框架,通过明确的失真定位来改进语音质量评估,解决了现有方法仅依赖于整体评分而缺乏局部失真信息的问题。
Existing text-to-image diffusion models often disregard physical constraints in projection mapping, leading to generated content that misaligns with real 3D scenes. To address this, this work proposes ConPhyG, a framework that formalizes two complementary generation paradigms: a cooperative mode that guides the diffusion process using pixel-level physical priors—such as depth, edges, and color gamut—and an adversarial mode that employs bounded numerical optimization to achieve radiometric compensation across multiple projectors, with dynamic switching between modes to balance artistic freedom and physical plausibility. ConPhyG further introduces a sequential generation strategy to ensure 360-degree multi-view consistent, physics-aware image synthesis. Experiments on a real four-projector system demonstrate that ConPhyG significantly outperforms existing methods, achieving state-of-the-art performance in geometric alignment, color gamut utilization, and semantic fidelity.
本文提出Ouroboros框架,利用语音增强模型的理想纯净输出作为自然触发器,在不注入外部触发器的情况下实现后门攻击,适用于多种模型和数据集。
本文提出GarmentWeaver,通过激活相关结构分支和采用预训练的视觉-语言模型来生成更准确且可执行的缝纫图样,解决了现有方法中结构与细节参数混淆的问题。
研究通过虚拟现实实验探讨了自动驾驶卡车与行人在多车道道路上的互动风险,并测试了三种干预措施,发现投影式外部人机界面最有效。
本文提出DAMOS框架,通过明确的失真定位来改进语音质量评估,解决了现有方法仅依赖于整体评分而缺乏局部失真信息的问题。
Existing text-to-image diffusion models often disregard physical constraints in projection mapping, leading to generated content that misaligns with real 3D scenes. To address this, this work proposes ConPhyG, a framework that formalizes two complementary generation paradigms: a cooperative mode that guides the diffusion process using pixel-level physical priors—such as depth, edges, and color gamut—and an adversarial mode that employs bounded numerical optimization to achieve radiometric compensation across multiple projectors, with dynamic switching between modes to balance artistic freedom and physical plausibility. ConPhyG further introduces a sequential generation strategy to ensure 360-degree multi-view consistent, physics-aware image synthesis. Experiments on a real four-projector system demonstrate that ConPhyG significantly outperforms existing methods, achieving state-of-the-art performance in geometric alignment, color gamut utilization, and semantic fidelity.