How I learned to stop worrying and love StopGrads: Stationarity, Convergence, and a case study on Flow Map Learning
本文解决了StopGrads在机器学习模型训练中影响收敛性的问题,通过引入一种回归原则来统一并优化StopGrad目标,保证了理论基础和训练效率。
本文解决了StopGrads在机器学习模型训练中影响收敛性的问题,通过引入一种回归原则来统一并优化StopGrad目标,保证了理论基础和训练效率。
本文探讨了通过合成任务扩展训练大型语言模型以设计小分子的方法,解决了直接使用成本高昂的化学评分函数进行在线训练的问题。
研究探讨了前沿大语言模型在连续和离散设置下的批量优化性能,发现其在语义丰富场景中表现更优。
This work addresses differentiable optimization on low-dimensional data manifolds embedded in high-dimensional spaces, where conventional gradient descent often deviates from the manifold and struggles with non-convex loss landscapes. The authors propose a novel approach that leverages diffusion and flow models to construct a diffeomorphic mapping, pulling the manifold back to a simple base space for optimization. Using tools from differential geometry, they prove this procedure is equivalent to Riemannian gradient descent, inherently preserving trajectories on the manifold. This is the first integration of diffeomorphic mappings with Riemannian optimization, extended to the Lie groups SO(3) and SE(3), yielding an automatic differentiation–compatible SO(3) gradient and a generalized adjoint-state backpropagation for Lie group ODE solvers. In protein design tasks, FrameFlow achieves a 91.3% secondary structure targeting success rate (versus 63.3% baseline), doubles the peptide binding affinity optimization speed compared to OC-Flow, and significantly reduces Rosetta energy by thousands of units.
This study addresses the challenge of bridging molecular-level protein–protein interactions with higher-order functional processes in disease through a cross-scale, interpretable modeling framework. The authors propose a hierarchical graph neural network that uniquely integrates the STRING protein–protein interaction network with the Reactome pathway hierarchy. Leveraging graph attention mechanisms, the model aggregates patient-specific multi-omics data—RNA-seq and DNA meth日晚间—bottom-up into coherent functional programs. A multi-layer supervised learning strategy enables interpretable integration from gene-level signals to biologically meaningful modules. Evaluated across ten cancer types in TCGA, the approach achieves over 90% accuracy, outperforming PPI-only models by 6.7% and surpassing single-head prediction by 12.3%. The method successfully recapitulates known oncogenic modules such as TP53–AKT and uncovers novel functional programs.
本文解决了StopGrads在机器学习模型训练中影响收敛性的问题,通过引入一种回归原则来统一并优化StopGrad目标,保证了理论基础和训练效率。
本文探讨了通过合成任务扩展训练大型语言模型以设计小分子的方法,解决了直接使用成本高昂的化学评分函数进行在线训练的问题。
研究探讨了前沿大语言模型在连续和离散设置下的批量优化性能,发现其在语义丰富场景中表现更优。
This work addresses differentiable optimization on low-dimensional data manifolds embedded in high-dimensional spaces, where conventional gradient descent often deviates from the manifold and struggles with non-convex loss landscapes. The authors propose a novel approach that leverages diffusion and flow models to construct a diffeomorphic mapping, pulling the manifold back to a simple base space for optimization. Using tools from differential geometry, they prove this procedure is equivalent to Riemannian gradient descent, inherently preserving trajectories on the manifold. This is the first integration of diffeomorphic mappings with Riemannian optimization, extended to the Lie groups SO(3) and SE(3), yielding an automatic differentiation–compatible SO(3) gradient and a generalized adjoint-state backpropagation for Lie group ODE solvers. In protein design tasks, FrameFlow achieves a 91.3% secondary structure targeting success rate (versus 63.3% baseline), doubles the peptide binding affinity optimization speed compared to OC-Flow, and significantly reduces Rosetta energy by thousands of units.
This study addresses the challenge of bridging molecular-level protein–protein interactions with higher-order functional processes in disease through a cross-scale, interpretable modeling framework. The authors propose a hierarchical graph neural network that uniquely integrates the STRING protein–protein interaction network with the Reactome pathway hierarchy. Leveraging graph attention mechanisms, the model aggregates patient-specific multi-omics data—RNA-seq and DNA meth日晚间—bottom-up into coherent functional programs. A multi-layer supervised learning strategy enables interpretable integration from gene-level signals to biologically meaningful modules. Evaluated across ten cancer types in TCGA, the approach achieves over 90% accuracy, outperforming PPI-only models by 6.7% and surpassing single-head prediction by 12.3%. The method successfully recapitulates known oncogenic modules such as TP53–AKT and uncovers novel functional programs.