Institution profile

Boston College

Academic institutionnorthamerica · us
Official website
Research library63linked papers
Opportunities0open roles
Selected work

Representative Papers

Trading off Consistency and Dimensionality of Convex Surrogates for the Mode

Feb 16, 2024arXiv.org

To address the intractability of surrogate loss optimization in large-scale multiclass classification caused by high-dimensional embeddings, this paper proposes a low-dimensional convex polyhedral embedding framework, establishing theoretical trade-offs among embedding dimension, consistency regions, and data distribution assumptions. First, it rigorously proves that “hallucination”—i.e., spurious class predictions—necessarily occurs when the embedding dimension is less than $n-1$. Second, under a low-noise assumption, it derives a verifiable consistency criterion. Third, it designs structured embeddings—including hypercubes and permutahedra—that achieve dimensionality reductions from $2^d$ to $d$ and from $d!$ to $d$, respectively. Finally, in the multiple-instance learning setting, it shows that full simplex consistency is guaranteed with only $n/2$ embedding dimensions and proves the existence of consistent subsets around any point-mass distribution.

1 citationsRead paper

A General Theory of Liquidity Provisioning for Automated Market Makers

Nov 15, 2023arXiv.org

This paper addresses the lack of a unified theoretical framework for liquidity provision in multi-asset automated market makers (AMMs) and prediction markets. Methodologically, it introduces the first general liquidity provision model based on a parallel market maker architecture, integrating game theory, convex analysis, and mechanism design to formalize multiple liquidity providers as cooperating, parallel market-making units that jointly quote prices and share risk. Contributions include: (1) the first rigorous proof that mainstream protocols—including Uniswap V2 and V3—are special cases of this framework; (2) a novel constrained structure and fee-allocation mechanism provably superior to Uniswap V3’s; and (3) a principled extension to prediction markets over arbitrary finite security sets and high-dimensional multi-asset AMMs, yielding a provably consistent and scalable design paradigm.

1 citationsRead paper

Algebraic Decomposition Theory for Transformer Length Generalization

Aug 13, 2026

This work addresses the lack of a precise characterization of Transformers’ length generalization capability on regular languages. By extending classical finite semigroup decomposition theory to the additive group of integers, the authors construct an algebraic framework based on the C-RASP formal system, thereby providing the first complete characterization of the class of regular languages amenable to length generalization by Transformers. This approach overcomes limitations of Krohn–Rhodes theory through the introduction of iterated cyclic decompositions and a polynomial-time decision algorithm. The resulting method not only efficiently determines whether any given regular language supports length generalization but also demonstrates significantly higher prediction accuracy than existing classification approaches in empirical evaluations.

0 citationsRead paper

Learning about Treatment Effects in Panels under Unknown Interference

Aug 13, 2026

This study addresses the challenge of disentangling treatment effects from unknown interference—such as spillovers—in panel data settings. The authors propose an identification strategy that does not require pre-specifying an exposure mapping or classifying affected units. By imposing pre-treatment fit constraints to bound donor weights and incorporating application-specific linear restrictions, they construct a sharp identified set for the treatment effect. Their approach integrates convex combination scaling, linear constraint modeling, bootstrap calibration, and confidence set inversion, enabling compatibility testing under only general assumptions. An empirical application to Arizona’s Legal Workers Act yields a 95% confidence identified set containing both positive and negative values, indicating that the sign of the treatment effect cannot be determined with the available data.

0 citationsRead paper

Stationary Errors and Quantile Regression in Short Panels

Aug 09, 2026

This study addresses the identification and estimation of common slope coefficients in quantile regression models with short panel data, accommodating unrestricted individual effects and temporally stationary disturbances. Under a stationarity assumption on the error term, the authors propose a two-step minimum distance estimator for fixed time dimension \(T\): first, a differencing identification strategy is constructed via intertemporal quantile regression projections, circumventing explicit estimation of individual effects; second, this restriction is leveraged to identify slope coefficients that are invariant across quantiles. The resulting estimator accommodates arbitrary within-individual serial correlation, achieves \(\sqrt{n}\)-consistency and asymptotic normality, and is accompanied by an analytical covariance estimator, a cluster-robust bootstrap procedure, and an overidentification test. Monte Carlo simulations demonstrate its excellent finite-sample performance across a range of data-generating processes.

0 citationsRead paper
Recent publications

Latest Papers

Algebraic Decomposition Theory for Transformer Length Generalization

Aug 13, 2026

This work addresses the lack of a precise characterization of Transformers’ length generalization capability on regular languages. By extending classical finite semigroup decomposition theory to the additive group of integers, the authors construct an algebraic framework based on the C-RASP formal system, thereby providing the first complete characterization of the class of regular languages amenable to length generalization by Transformers. This approach overcomes limitations of Krohn–Rhodes theory through the introduction of iterated cyclic decompositions and a polynomial-time decision algorithm. The resulting method not only efficiently determines whether any given regular language supports length generalization but also demonstrates significantly higher prediction accuracy than existing classification approaches in empirical evaluations.

0 citationsRead paper

Learning about Treatment Effects in Panels under Unknown Interference

Aug 13, 2026

This study addresses the challenge of disentangling treatment effects from unknown interference—such as spillovers—in panel data settings. The authors propose an identification strategy that does not require pre-specifying an exposure mapping or classifying affected units. By imposing pre-treatment fit constraints to bound donor weights and incorporating application-specific linear restrictions, they construct a sharp identified set for the treatment effect. Their approach integrates convex combination scaling, linear constraint modeling, bootstrap calibration, and confidence set inversion, enabling compatibility testing under only general assumptions. An empirical application to Arizona’s Legal Workers Act yields a 95% confidence identified set containing both positive and negative values, indicating that the sign of the treatment effect cannot be determined with the available data.

0 citationsRead paper

Stationary Errors and Quantile Regression in Short Panels

Aug 09, 2026

This study addresses the identification and estimation of common slope coefficients in quantile regression models with short panel data, accommodating unrestricted individual effects and temporally stationary disturbances. Under a stationarity assumption on the error term, the authors propose a two-step minimum distance estimator for fixed time dimension \(T\): first, a differencing identification strategy is constructed via intertemporal quantile regression projections, circumventing explicit estimation of individual effects; second, this restriction is leveraged to identify slope coefficients that are invariant across quantiles. The resulting estimator accommodates arbitrary within-individual serial correlation, achieves \(\sqrt{n}\)-consistency and asymptotic normality, and is accompanied by an analytical covariance estimator, a cluster-robust bootstrap procedure, and an overidentification test. Monte Carlo simulations demonstrate its excellent finite-sample performance across a range of data-generating processes.

0 citationsRead paper

Clustered Attractor Manifolds and Dynamical Condensation in Self-Attention

Aug 09, 2026

This work investigates how state-dependent dynamic feedback in self-attention mechanisms spontaneously gives rise to structured representations and their associated attractor geometries. By constructing a simplified normalized self-attention dynamical model and leveraging tools from statistical physics—specifically the thermodynamic limit, high-dimensional geometry, and dynamical systems theory—the study elucidates the interplay between token clustering and attention matrix formation. The authors introduce the “overlap gap” as a key order parameter governing attractor structure and identify a critical threshold in attention sharpness: only when this threshold is exceeded and intra-cluster similarity substantially surpasses inter-cluster similarity does the system undergo a dynamic condensation phase transition from an unstructured initial state, thereby self-organizing into stable clustered attractor manifolds. In this regime, cross-cluster attention decays exponentially with dimensionality, yielding a continuous spectrum of structures ranging from macroscopic clusters to microscopic fragments.

0 citationsRead paper

Classical $\mathrm{SU}(2)$ Models Match or Exceed Shallow Variational Quantum Circuits on Vision Benchmarks

Aug 07, 2026

This study systematically evaluates the performance of classical and quantum models sharing an $\mathrm{SU}(2)$ geometric structure on visual recognition tasks, investigating whether shallow variational quantum circuits offer practical advantages. Building upon frozen features from a pretrained ResNet18 backbone, the authors compare real-valued, quaternion-based, and variational quantum classifiers—with and without entanglement—across MNIST, FashionMNIST, and CIFAR-10, optimizing all models using Fubini–Study natural gradients. The results demonstrate that quaternion networks match or closely approach real-valued baselines while significantly outperforming quantum counterparts; entanglement yields only marginal gains on grayscale images and degrades performance when applied to pretrained features. This work provides the first evidence that merely sharing an $\mathrm{SU}(2)$ structure is insufficient to confer a quantum advantage in such settings.

0 citationsRead paper