Fenchel-Young Duality Gaps: Certified Early Stopping for Regularized Inverse Problems
本文研究了正则化逆问题中的可计算误差界和认证的提前停止方法,通过Fenchel-Young对偶间隙将总误差分解为数据保真度损失和正则化损失,提供了一种基于对偶间隙下降的提前停止规则。
本文研究了正则化逆问题中的可计算误差界和认证的提前停止方法,通过Fenchel-Young对偶间隙将总误差分解为数据保真度损失和正则化损失,提供了一种基于对偶间隙下降的提前停止规则。
本文提出Debias-SparseGPT,一种结合了代表性偏差减少的后训练剪枝方法,有效解决了现有稀疏化方法在大语言模型中放大偏见的问题。
为解决多指抓取对新物体泛化能力差的问题,提出GOAG模型,通过学习夹爪接触面分布的紧凑潜在表示来生成有效的抓取配置。
本文提出CoToGrasp框架,通过学习规范工作空间来生成稳定且多样的抓握方式,解决了现有灵巧抓握规划器仅优化物理稳定性的问题。
This work addresses the significant drop in robustness of existing AI-generated music detectors when confronted with simple audio transformations such as tempo changes and pitch shifts. To overcome this limitation, the authors propose a novel detection architecture that inherently incorporates frequency scaling invariance. The approach maps audio signals onto a logarithmic frequency axis via log-STFT and combines learnable cross-correlation filters with max-pooling to achieve translation invariance during inference. This is the first method to integrate frequency scaling invariance directly into the detection pipeline, simultaneously producing both a binary authenticity decision and an estimate of the applied tempo scaling factor, thereby enhancing model interpretability and adversarial robustness. Experimental results demonstrate that the proposed method maintains high detection accuracy under various audio transformation attacks, substantially outperforming current state-of-the-art techniques while accurately estimating transformation parameters.
本文研究了正则化逆问题中的可计算误差界和认证的提前停止方法,通过Fenchel-Young对偶间隙将总误差分解为数据保真度损失和正则化损失,提供了一种基于对偶间隙下降的提前停止规则。
本文提出Debias-SparseGPT,一种结合了代表性偏差减少的后训练剪枝方法,有效解决了现有稀疏化方法在大语言模型中放大偏见的问题。
为解决多指抓取对新物体泛化能力差的问题,提出GOAG模型,通过学习夹爪接触面分布的紧凑潜在表示来生成有效的抓取配置。
本文提出CoToGrasp框架,通过学习规范工作空间来生成稳定且多样的抓握方式,解决了现有灵巧抓握规划器仅优化物理稳定性的问题。
This work addresses the significant drop in robustness of existing AI-generated music detectors when confronted with simple audio transformations such as tempo changes and pitch shifts. To overcome this limitation, the authors propose a novel detection architecture that inherently incorporates frequency scaling invariance. The approach maps audio signals onto a logarithmic frequency axis via log-STFT and combines learnable cross-correlation filters with max-pooling to achieve translation invariance during inference. This is the first method to integrate frequency scaling invariance directly into the detection pipeline, simultaneously producing both a binary authenticity decision and an estimate of the applied tempo scaling factor, thereby enhancing model interpretability and adversarial robustness. Experimental results demonstrate that the proposed method maintains high detection accuracy under various audio transformation attacks, substantially outperforming current state-of-the-art techniques while accurately estimating transformation parameters.