Institution profile

Simons Foundation

Academic institutionnorthamerica · us
Official website
Research library18linked papers
Opportunities0open roles
Selected work

Representative Papers

A Theory of Initialisation's Impact on Specialisation

Mar 04, 2025

This work challenges the necessity of neuron specialization for mitigating catastrophic forgetting in continual learning, revealing that specialization is primarily governed by network initialization rather than intrinsic task properties. Method: Through theoretical analysis and empirical validation, the authors demonstrate that weight imbalance and high weight entropy actively induce localized representations; they provide the first theoretical proof that specialization is not inherent but contingent. They further derive a quantitative relationship between specialization degree and initialization parameters, and reproduce the monotonic relationship between task similarity and forgetting rate even in non-specialized networks. Contribution/Results: Specialized initialization significantly enhances Elastic Weight Consolidation (EWC) performance, an effect attributable to initialization-induced prior shaping of representation structure. These findings establish a novel theoretical foundation for regularization design in continual learning and yield principled guidelines for initialization strategy selection.

1 citationsRead paper

GEOPHYS: The Geometry of Physical Plausibility

Jun 15, 2026

This work addresses the computational expense, reliance on external models, or need for training modifications that plague existing methods for evaluating physical plausibility in videos. The authors propose GEOPHYS, which reveals—for the first time—that frame-wise embeddings from a frozen image encoder inherently encode five geometric properties that serve as effective signals of physical reasonableness. Leveraging this insight, GEOPHYS efficiently discriminates physically implausible events without any additional training and functions as a best-of-N verifier for physics alignment in video generation. Experiments demonstrate its superior performance, achieving 98.3% and 93.3% accuracy on LikePhys and IntPhys2 benchmarks, respectively. When applied to MAGI-1 24B video generation, it improves physical plausibility to 64.50% while reducing computational time by 1.5× and memory usage by 4.65× compared to prior approaches.

0 citationsRead paper

Flexible Online Representation Learning Based on Similarity Matching

May 31, 2026

This work addresses the challenge of online learning of high-dimensional sparse representations under row-sum constraints—such as doubly stochasticity—in large-scale settings. The authors propose a similarity-matching-based online learning algorithm that circumvents computationally expensive relaxations like completely positive or doubly nonnegative matrix factorizations. By incorporating translation invariance and explicit sparsity constraints, the method flexibly accommodates diverse tasks including clustering, manifold tiling, and sparse coding. Notably, it achieves the first efficient online sparse representation learning framework that enforces row-sum constraints, thereby significantly enhancing scalability and practicality for large-scale applications while preserving biological interpretability.

0 citationsRead paper

Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics

May 28, 2026

This work investigates how attention mechanisms can perform effective Bayesian inference and denoising under full-token corruption. We interpret single-layer attention as a kernel-weighted posterior mean estimator based on the empirical distribution of contextual tokens, and characterize the progressive refinement of this empirical distribution in deep networks through particle dynamics. By incorporating long-range skip connections, the architecture realizes a two-stage inference process. Theoretical analysis demonstrates that, under fixed kernel bandwidth and finite integration time, the method achieves effective denoising without explicit noise scheduling, and the empirical estimator asymptotically converges to the Bayes-optimal predictor. This study elucidates the distinct roles of network depth and attention residuals in statistical inference and provides theoretical guarantees for posterior mean recovery.

0 citationsRead paper
Recent publications

Latest Papers

GEOPHYS: The Geometry of Physical Plausibility

Jun 15, 2026

This work addresses the computational expense, reliance on external models, or need for training modifications that plague existing methods for evaluating physical plausibility in videos. The authors propose GEOPHYS, which reveals—for the first time—that frame-wise embeddings from a frozen image encoder inherently encode five geometric properties that serve as effective signals of physical reasonableness. Leveraging this insight, GEOPHYS efficiently discriminates physically implausible events without any additional training and functions as a best-of-N verifier for physics alignment in video generation. Experiments demonstrate its superior performance, achieving 98.3% and 93.3% accuracy on LikePhys and IntPhys2 benchmarks, respectively. When applied to MAGI-1 24B video generation, it improves physical plausibility to 64.50% while reducing computational time by 1.5× and memory usage by 4.65× compared to prior approaches.

0 citationsRead paper

Flexible Online Representation Learning Based on Similarity Matching

May 31, 2026

This work addresses the challenge of online learning of high-dimensional sparse representations under row-sum constraints—such as doubly stochasticity—in large-scale settings. The authors propose a similarity-matching-based online learning algorithm that circumvents computationally expensive relaxations like completely positive or doubly nonnegative matrix factorizations. By incorporating translation invariance and explicit sparsity constraints, the method flexibly accommodates diverse tasks including clustering, manifold tiling, and sparse coding. Notably, it achieves the first efficient online sparse representation learning framework that enforces row-sum constraints, thereby significantly enhancing scalability and practicality for large-scale applications while preserving biological interpretability.

0 citationsRead paper

Attention as In-Context Empirical Bayes: A Two-Stage View via Particle Dynamics

May 28, 2026

This work investigates how attention mechanisms can perform effective Bayesian inference and denoising under full-token corruption. We interpret single-layer attention as a kernel-weighted posterior mean estimator based on the empirical distribution of contextual tokens, and characterize the progressive refinement of this empirical distribution in deep networks through particle dynamics. By incorporating long-range skip connections, the architecture realizes a two-stage inference process. Theoretical analysis demonstrates that, under fixed kernel bandwidth and finite integration time, the method achieves effective denoising without explicit noise scheduling, and the empirical estimator asymptotically converges to the Bayes-optimal predictor. This study elucidates the distinct roles of network depth and attention residuals in statistical inference and provides theoretical guarantees for posterior mean recovery.

0 citationsRead paper

Bifunction and Interlevel Delaunay Trifiltrations

May 20, 2026

This work proposes the first three-parameter Delaunay trifiltration for point clouds equipped with ℝ²-valued functions, satisfying weak topological equivalence to enable multiparameter persistent homology in time-varying data. Extending the classical Delaunay filtration to higher-dimensional function-valued settings, the method preserves weak equivalence with the offset filtration while introducing an efficient algorithm with time complexity O(|X|^{⌈d/2⌉+2}) and a computational framework whose memory usage grows nearly linearly with input size. Experimental results demonstrate that the approach effectively handles thousands of points in ℝ³, offering both scalability and practicality. This framework thus provides a computationally feasible new tool for multiparameter topological data analysis.

0 citationsRead paper