Equivariance Breaks the Learning Rate
研究解决了等变网络中Adam优化器性能不佳的问题,通过单独归一化每个块更新的方法,提高了其在等变线性层中的表现。
研究解决了等变网络中Adam优化器性能不佳的问题,通过单独归一化每个块更新的方法,提高了其在等变线性层中的表现。
本文探讨了使用Hyper-V套接字作为恶意软件分析沙箱的实时数据提取通道,对比了其与WinSock TCP套接字在阻塞和枚举方面的优势,并比较了两者的数据吞吐量。
While large foundation models demonstrate strong performance in solving time-dependent partial differential equations, their high computational cost limits their practicality as replacements for efficient numerical solvers. This work proposes the Teacher Rollout Extension (TREX) framework, which leverages knowledge distillation to transfer capabilities from a pretrained teacher model to a lightweight student model. By using long-horizon synthetic trajectories generated by the teacher to augment limited downstream data, TREX enables sampling of rollout trajectories without requiring prior knowledge of the initial condition distribution. This exposes the student model to both long-term dynamics and local recovery behaviors, while allowing integration of task-specific inductive biases—such as equivariance. Combined with noise injection and an equivariant network architecture, the resulting student model achieves several orders of magnitude fewer parameters, over tenfold faster inference, and accuracy comparable to or exceeding that of the teacher.
This work investigates the intrinsic mechanisms by which foundation models detect deepfakes, addressing why pretrained representations effectively distinguish authentic from synthetic media. The study reveals that forged samples consistently elicit lower-magnitude feature responses across diverse foundation models and systematically demonstrates— for the first time—that this amplitude discrepancy serves as a key signal for authenticity verification, rooted in semantic shift. Building on this insight, the authors reformulate deepfake detection as an anomaly detection task, showing that simple statistics of feature magnitudes alone enable efficient zero-shot detection. The proposed approach achieves performance on par with complex specialized models across both image and video modalities, with detection capability scaling favorably with model size, thereby confirming that large-scale foundation models inherently possess strong zero-shot potential for forgery identification.
This work addresses catastrophic forgetting in parameter-efficient continual learning by proposing TailLoR, a method that constructs a fixed reference frame based on the singular vectors of pre-trained weights and applies low-rank updates to the singular value matrix. TailLoR introduces a soft spectral penalty to steer adaptation away from dominant singular directions, thereby channeling parameter updates toward the long-tail spectral coordinates. This approach is the first to explicitly preserve principal components within a spectral decomposition framework while leveraging the long-tail spectrum for flexible adaptation, effectively balancing model stability and plasticity. Experimental results demonstrate that TailLoR substantially reduces interference across tasks, significantly improves continual learning performance, and achieves these gains with an extremely small number of trainable parameters.
研究解决了等变网络中Adam优化器性能不佳的问题,通过单独归一化每个块更新的方法,提高了其在等变线性层中的表现。
本文探讨了使用Hyper-V套接字作为恶意软件分析沙箱的实时数据提取通道,对比了其与WinSock TCP套接字在阻塞和枚举方面的优势,并比较了两者的数据吞吐量。
While large foundation models demonstrate strong performance in solving time-dependent partial differential equations, their high computational cost limits their practicality as replacements for efficient numerical solvers. This work proposes the Teacher Rollout Extension (TREX) framework, which leverages knowledge distillation to transfer capabilities from a pretrained teacher model to a lightweight student model. By using long-horizon synthetic trajectories generated by the teacher to augment limited downstream data, TREX enables sampling of rollout trajectories without requiring prior knowledge of the initial condition distribution. This exposes the student model to both long-term dynamics and local recovery behaviors, while allowing integration of task-specific inductive biases—such as equivariance. Combined with noise injection and an equivariant network architecture, the resulting student model achieves several orders of magnitude fewer parameters, over tenfold faster inference, and accuracy comparable to or exceeding that of the teacher.
This work investigates the intrinsic mechanisms by which foundation models detect deepfakes, addressing why pretrained representations effectively distinguish authentic from synthetic media. The study reveals that forged samples consistently elicit lower-magnitude feature responses across diverse foundation models and systematically demonstrates— for the first time—that this amplitude discrepancy serves as a key signal for authenticity verification, rooted in semantic shift. Building on this insight, the authors reformulate deepfake detection as an anomaly detection task, showing that simple statistics of feature magnitudes alone enable efficient zero-shot detection. The proposed approach achieves performance on par with complex specialized models across both image and video modalities, with detection capability scaling favorably with model size, thereby confirming that large-scale foundation models inherently possess strong zero-shot potential for forgery identification.
This work addresses catastrophic forgetting in parameter-efficient continual learning by proposing TailLoR, a method that constructs a fixed reference frame based on the singular vectors of pre-trained weights and applies low-rank updates to the singular value matrix. TailLoR introduces a soft spectral penalty to steer adaptation away from dominant singular directions, thereby channeling parameter updates toward the long-tail spectral coordinates. This approach is the first to explicitly preserve principal components within a spectral decomposition framework while leveraging the long-tail spectrum for flexible adaptation, effectively balancing model stability and plasticity. Experimental results demonstrate that TailLoR substantially reduces interference across tasks, significantly improves continual learning performance, and achieves these gains with an extremely small number of trainable parameters.