AURA: Unified Multimodal Framework for Conversational Music Editing
AURA通过多模态大语言模型和概念到音频模块解决音乐编辑中逐步精细化的问题,实现精确编辑同时保留未受影响内容。
AURA通过多模态大语言模型和概念到音频模块解决音乐编辑中逐步精细化的问题,实现精确编辑同时保留未受影响内容。
研究使用VGG19、MobileNetV2等模型及GAN生成的数据增强技术解决胸部X光片中的肺炎检测问题,但未证明GAN增强方法的有效性。
本文通过联邦学习和低秩适应(LoRA)方法,在不交换数据的情况下,改进了跨四个国际胸部X光队列的BiomedCLIP模型的性能。
This study addresses the suboptimal training caused by fixed surrogate gradients in spiking Transformers by proposing the SAGE mechanism. SAGE quantifies uncertainty via normalized self-attention entropy to adaptively modulate surrogate gradient slopes, enabling dynamic parameter optimization with zero inference overhead. By preserving the inference model unchanged, this approach significantly enhances optimization flexibility. Experiments on CIFAR-10/100 demonstrate that SAGE consistently improves accuracy by 1–2% over fixed baselines, effectively resolving the trade-off between gradient estimation and model performance in spiking neural networks.
This study addresses the lack of standardized implementation of Grad-CAM in Vision Transformers (ViTs), which has led to ambiguous formulations, poor reproducibility, and inconsistent interpretations. Through a systematic review of 175 relevant works, this paper proposes the first descriptive taxonomy for Grad-CAM variants tailored to ViTs, clarifying implicit assumptions and inconsistencies across critical components—namely feature token selection, gradient target specification, spatial reconstruction, and aggregation strategies. The analysis reveals that most existing studies inadequately document implementation details, demonstrating that the adaptation of Grad-CAM to ViTs is far from a trivial extension of its CNN-based counterpart. By establishing a clear and rigorous methodological framework, this work aims to enhance transparency, reproducibility, and reliability in explainable artificial intelligence for transformer-based vision models.
AURA通过多模态大语言模型和概念到音频模块解决音乐编辑中逐步精细化的问题,实现精确编辑同时保留未受影响内容。
研究使用VGG19、MobileNetV2等模型及GAN生成的数据增强技术解决胸部X光片中的肺炎检测问题,但未证明GAN增强方法的有效性。
本文通过联邦学习和低秩适应(LoRA)方法,在不交换数据的情况下,改进了跨四个国际胸部X光队列的BiomedCLIP模型的性能。
This study addresses the suboptimal training caused by fixed surrogate gradients in spiking Transformers by proposing the SAGE mechanism. SAGE quantifies uncertainty via normalized self-attention entropy to adaptively modulate surrogate gradient slopes, enabling dynamic parameter optimization with zero inference overhead. By preserving the inference model unchanged, this approach significantly enhances optimization flexibility. Experiments on CIFAR-10/100 demonstrate that SAGE consistently improves accuracy by 1–2% over fixed baselines, effectively resolving the trade-off between gradient estimation and model performance in spiking neural networks.
This study addresses the lack of standardized implementation of Grad-CAM in Vision Transformers (ViTs), which has led to ambiguous formulations, poor reproducibility, and inconsistent interpretations. Through a systematic review of 175 relevant works, this paper proposes the first descriptive taxonomy for Grad-CAM variants tailored to ViTs, clarifying implicit assumptions and inconsistencies across critical components—namely feature token selection, gradient target specification, spatial reconstruction, and aggregation strategies. The analysis reveals that most existing studies inadequately document implementation details, demonstrating that the adaptation of Grad-CAM to ViTs is far from a trivial extension of its CNN-based counterpart. By establishing a clear and rigorous methodological framework, this work aims to enhance transparency, reproducibility, and reliability in explainable artificial intelligence for transformer-based vision models.