Institution profile

McMaster University

Academic institutionnorthamerica · ca
Official website
Research library262linked papers
Opportunities0open roles
Selected work

Representative Papers

Mixture of Experts Softens the Curse of Dimensionality in Operator Learning

Apr 13, 2024

To address the computational and memory bottlenecks imposed by the curse of dimensionality in high-dimensional operator learning, this paper proposes the Mixture-of-Experts Neural Operator (MoNO): a framework that decomposes a global nonlinear operator into multiple lightweight expert sub-operators, routed input-adaptively via a learnable decision-tree mechanism. Theoretically, we establish the first distributed universal approximation theorem, proving that MoNO uniformly approximates any Lipschitz-continuous nonlinear operator in Sobolev spaces; each expert’s depth, width, and rank scale as O(ε⁻¹), ensuring controllable memory footprint compatible with standard hardware. We further derive the first quantitative approximation rate for classical neural operators. Experiments and theory jointly demonstrate that MoNO achieves ε-accuracy with significantly reduced complexity, overcoming both expressive and deployability limitations inherent to monolithic neural operators.

20 citationsRead paper

Extremal cases of distortion risk measures with partial information

Apr 21, 2024

This paper addresses the extremal problem of generalized distorted risk measures under weak distributional information—specifically, when only the first two moments and shape priors (e.g., symmetry or unimodality) are known. We develop the first unified analytical framework integrating probabilistic inequalities, extreme distribution theory, and convex order optimization. This enables the derivation of explicit, tight upper and lower bounds for distorted risk measures subject to shape constraints—results previously unavailable in the literature. Unlike conventional approaches requiring full distributional specification, our method eliminates model-class dependence, thereby substantially enhancing robustness in risk assessment under model uncertainty—particularly in finance and insurance. The work establishes a rigorous theoretical foundation and provides practical computational tools for robust risk measurement under partial information.

2 citationsRead paper

A Comparative Study of Technical Writing Feedback Quality: Evaluating LLMs, SLMs, and Humans in Computer Science Topics

Dec 01, 2025arXiv.org

This study addresses the scalability challenges of providing technical writing feedback in computer science education by proposing a novel AI-assisted pedagogical paradigm that ensures privacy preservation and incurs zero marginal cost. Employing a mixed-methods approach, the research compares the quality of feedback generated by a locally deployed, quantized small language model (SLM) based on Llama-3.1, commercial large language models (e.g., GPT-4), and human instructors across programming, operating systems, and writing seminar courses. Empirical results indicate that the SLM is preferred by students for its readability and actionable suggestions, delivering feedback quality comparable to or exceeding that of commercial LLMs, while human instructors retain an advantage in highly specialized tasks. The findings validate the efficacy of a tiered collaboration model wherein AI provides structured guidance and instructors focus on higher-order conceptual instruction.

1 citationsRead paper

Retinex-guided Histogram Transformer for Mask-free Shadow Removal

Apr 18, 2025

Existing shadow removal methods often rely on hard-to-obtain shadow masks, limiting their generalizability. To address this, we propose ReHiT—the first mask-free single-image shadow removal framework grounded in Retinex theory. ReHiT employs a dual-branch CNN-Transformer architecture to separately model reflectance and illumination. We introduce the Illumination-Guided Histogram Transformer Block (IGHB), the first of its kind, which integrates Retinex-based decomposition principles with multi-scale semantic modeling to accurately capture complex, spatially varying shadows under non-uniform illumination. Additionally, residual dense feature learning is incorporated to enhance representational capacity. Evaluated on the NTIRE 2025 benchmark, ReHiT achieves state-of-the-art performance with the smallest parameter count and fastest inference speed, significantly improving practical deployability in real-world scenarios.

1 citationsRead paper

Deep Kalman Filters Can Filter

Oct 30, 2023Social Science Research Network

Conventional deep Kalman filters (DKFs) lack theoretical guarantees for non-Markovian, conditionally Gaussian signal processes, limiting their applicability in mathematically rigorous domains such as mathematical finance. Method: We propose the continuous-time deep Kalman filter (CT-DKF), modeling signal dynamics via stochastic differential equations and quantifying distributional approximation error using the 2-Wasserstein distance. Our framework establishes uniform approximation of conditional distributions over regular compact path spaces. Contribution/Results: CT-DKF is the first DKF variant with a rigorous theoretical guarantee of consistent approximation to the optimal Bayesian filter for arbitrary non-Markovian, conditionally Gaussian processes. It overcomes the absence of convergence and generalization guarantees in existing DKFs; its approximation error is precisely bounded by the worst-case 2-Wasserstein distance. This provides a verifiable theoretical foundation for financial applications including bond and option pricing, and model calibration.

1 citationsRead paper
Recent publications

Latest Papers