Institution profile

Morgan Stanley

Industry researchnorthamerica · us
Official website
Research library61linked papers
Opportunities0open roles
Selected work

Representative Papers

Generalized Discrete Diffusion with Self-Correction

Feb 13, 2026

Existing self-correction methods for discrete diffusion models are typically introduced during inference or post-training, exhibiting limited generalization and often degrading performance. This work proposes the Self-Correcting Discrete Diffusion (SCDD) model, which, for the first time, integrates an explicit state-transition-based self-correction mechanism directly into pretraining within a discrete-time framework. SCDD eliminates redundant re-masking steps and relies solely on a uniform absorption objective for learning. By simplifying the noise schedule and combining BERT-style pretraining with parallel decoding, the method substantially improves decoding efficiency—demonstrated on GPT-2-scale experiments—while preserving high generation quality.

1 citationsRead paper

Elliptic Loss Regularization

Mar 04, 2025

This work addresses the weak extrapolation capability and unreliable predictions in unseen regions of neural networks under distribution shift and group imbalance. We propose a novel smoothness regularization method grounded in elliptic partial differential equation (PDE) theory. Specifically, the method constrains the second-order differential geometric structure of the loss function in input space via an elliptic operator, explicitly enforcing elliptic PDE properties to enhance local smoothness and extrapolation stability of the loss landscape. To our knowledge, this is the first systematic integration of elliptic PDE theory into deep learning regularization design. Theoretically, we derive a controllable upper bound on generalization error for unseen regions under this regularization. Empirically, it significantly improves robustness across multiple out-of-distribution generalization and fairness benchmarks, while remaining fully compatible with standard training pipelines and incurring minimal computational overhead.

1 citationsRead paper

Parabolic Continual Learning

Mar 03, 2025

To address the coupled generalization-forgetting error problem arising from catastrophic forgetting in continual learning, this paper pioneers modeling the temporal evolution of the loss function as a parabolic partial differential equation (PDE), with the memory buffer serving as a dynamic boundary condition—explicitly capturing long-range dependencies and error propagation. Leveraging the intrinsic physical regularity of PDEs, we formulate a spatiotemporal constrained optimization framework driven by boundary conditions, enabling analyzable and interpretable dynamic regularization. Theoretically, we derive a tight coupled bound on forgetting and generalization errors. Empirically, our method significantly reduces forgetting across multiple standard benchmarks, and the theoretical error bound closely aligns with observed performance—demonstrating the effectiveness, stability, and analytical tractability of PDE-based regularization.

1 citationsRead paper

Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models

Jul 17, 2026

Existing diffusion language models typically employ fixed-depth, single-step look-ahead decoding strategies, which struggle to balance efficiency and accuracy in long-horizon generation and fail to accommodate the heterogeneity of intermediate states. This work proposes AdaLook, a novel framework that introduces, for the first time, a dynamic multi-step look-ahead mechanism guided by the variance of candidate scores. AdaLook adaptively decides whether to further unfold or expand search branches, thereby avoiding unnecessary deep computations and enabling re-initiation of look-ahead from informative intermediate states. By integrating masked diffusion language modeling with adaptive decision-making and branch expansion strategies, AdaLook substantially outperforms existing single-step approaches across multiple benchmarks, achieving comparable generation quality with significantly fewer decoding steps.

0 citationsRead paper
Recent publications

Latest Papers

Adaptive Multi-Step Lookahead Decoding for Diffusion Language Models

Jul 17, 2026

Existing diffusion language models typically employ fixed-depth, single-step look-ahead decoding strategies, which struggle to balance efficiency and accuracy in long-horizon generation and fail to accommodate the heterogeneity of intermediate states. This work proposes AdaLook, a novel framework that introduces, for the first time, a dynamic multi-step look-ahead mechanism guided by the variance of candidate scores. AdaLook adaptively decides whether to further unfold or expand search branches, thereby avoiding unnecessary deep computations and enabling re-initiation of look-ahead from informative intermediate states. By integrating masked diffusion language modeling with adaptive decision-making and branch expansion strategies, AdaLook substantially outperforms existing single-step approaches across multiple benchmarks, achieving comparable generation quality with significantly fewer decoding steps.

0 citationsRead paper

A Knowledge Theory of Capital:The Value of Natural and Artificial Intelligence

Jun 12, 2026

This study addresses the limitations of traditional capital theory in explaining value creation by artificial and natural intelligence in knowledge-driven economies. It introduces the concept of “knowledge-embodied capital,” reconceptualizing software, data, models, and organizational structures as accumulable, governable, and reconfigurable knowledge assets. The framework distinguishes five forms of knowledge: embodied, disembodied, institutionalized, commons-based, and public. Integrating political economy, institutional analysis, and information science, the work develops an original conditional theory of knowledge governance, featuring novel constructs such as initial transformation, cognitive enclosure, and feedback capture. The research demonstrates that the core of contemporary wealth lies not merely in capital accumulation but fundamentally in the governance of productive knowledge, thereby offering a new paradigm for evaluating the value generated by both artificial and natural intelligence.

0 citationsRead paper

TimeRouter: Efficient and Adaptive Routing of Time-Series Foundation Models

Jun 09, 2026

This work addresses the inefficiency of existing time series foundation model (TSFM) pools, which lack adaptive expert selection mechanisms and incur high inference overhead due to reliance on large language model (LLM) controllers. To overcome this limitation, we propose TimeRouter—the first framework enabling LLM-free, adaptive routing among pretrained TSFMs. TimeRouter employs a lightweight discriminative routing head, selective gating, and an ensemble fallback mechanism to dynamically dispatch inputs to the most suitable models in the pool. This approach significantly enhances system modularity and inference efficiency, achieving state-of-the-art performance on the GIFT-EVAL benchmark with an LB MASE of 0.6765. Furthermore, our experiments demonstrate that both the composition of the model pool and the design of the gating mechanism critically influence routing effectiveness.

0 citationsRead paper

Learning to Strategically Acquire Resources in Competition

Jun 04, 2026

This study addresses strategic competition among multiple agents for divisible, scarce resources—such as financial assets or computational capacity—under conditions of asymmetric information or the absence of a common prior. By developing a game-theoretic framework that integrates market price dynamics, the work extends the existence, uniqueness, and computationally efficient characterization of Bayesian Nash equilibria to partial-information settings with a common prior, and further establishes convergence guarantees for simultaneous learning dynamics when no common prior is assumed. Theoretical contributions include a rigorous characterization of equilibrium properties and an upper bound on the Price of Anarchy. Empirically, simulations based on real-world financial data validate both the convergence of the proposed algorithms and the efficacy of the resulting strategic behaviors.

0 citationsRead paper