Institution profile

George Washington University

Academic institutionnorthamerica · us
Official website
Research library328linked papers
Opportunities0open roles
Selected work

Representative Papers

Structuring Value Representations via Geometric Coherence in Markov Decision Processes

Feb 03, 2026

This work addresses the challenges of instability and low sample efficiency in value function estimation within reinforcement learning by introducing, for the first time, an order-theoretic perspective. The authors formulate value learning as a partially ordered set (poset) learning problem and propose the GCR-RL framework, which progressively refines a hyper-poset structure guided by temporal difference signals to ensure geometric consistency in value representations. Building on this foundation, they develop two novel algorithms compatible with both Q-learning and Actor-Critic architectures, accompanied by theoretical convergence guarantees. Empirical evaluations demonstrate that the proposed approach significantly improves sample efficiency and training stability across a variety of tasks, outperforming several strong baselines.

1 citationsRead paper

Manifold-Constrained Energy-Based Transition Models for Offline Reinforcement Learning

Feb 02, 2026

This work addresses the challenges of value overestimation and policy degradation in offline reinforcement learning caused by distributional shift, particularly in data-sparse regions. To mitigate these issues, the authors propose the Manifold-Constrained Energy-based Transition Model (MC-ETM), which trains a conditional energy-based model in latent space by integrating manifold projection with diffusion-based negative sampling. The method sharpens the energy landscape by generating near-manifold hard negative samples via Langevin dynamics. Furthermore, an energy-guided truncation mechanism combined with pessimistic Bellman backups is introduced to establish a hybrid pessimistic MDP framework. Experimental results demonstrate that MC-ETM significantly improves multi-step dynamics fidelity and normalized returns on standard offline control benchmarks, outperforming existing approaches—especially in scenarios involving irregular dynamics and sparse data.

1 citationsRead paper

Geometry of Drifting MDPs with Path-Integral Stability Certificates

Jan 29, 2026

This work addresses the challenge of non-stationary environmental dynamics and rewards—such as drifts, oscillations, or abrupt policy shifts—in real-world reinforcement learning, which often induce policy jitter and tracking errors. Existing methods struggle to capture the local geometric structure of such non-stationarity. The authors model non-stationary discounted MDPs as differentiable homotopy paths and quantify the intrinsic complexity of environmental changes through path length, curvature, and inflection points along the trajectory of optimal Bellman fixed points. Leveraging this geometric characterization, they adaptively modulate learning and planning intensity. For the first time, stability bounds based on path integrals and a gap-aware safety-feasibility region are established from a geometric perspective, enabling formal certification of stability in policy-switching regions. The proposed lightweight algorithms, HT-RL and HT-MCTS, significantly outperform static baselines in oscillatory and high-switching scenarios, effectively reducing dynamic regret and improving policy tracking performance.

1 citationsRead paper

Immunological Density Shapes Recovery Trajectories in Long COVID

Jan 09, 2026

This study investigates the drivers of clinical recovery in Long COVID, disentangling the effects of natural disease progression from those of vaccination. Leveraging longitudinal data from 13,511 patients encompassing 97,564 clinical assessments and vaccination records, and applying a clinically validated PASC symptom threshold (≥12 symptoms), the research identifies three distinct recovery trajectories: Protected, Refractory, and Responders. The analysis reveals that symptom severity exhibits a slight upward trend over time, with spontaneous remission being rare. Crucially, cumulative vaccine doses are significantly and negatively associated with symptom burden, indicating that repeated immunization plays a pivotal role in promoting recovery. Furthermore, baseline symptom severity demonstrates strong predictive value for clinical outcomes, underscoring its utility as a prognostic indicator.

1 citationsRead paper

ACDZero: MCTS Agent for Mastering Automated Cyber Defense

Jan 05, 2026arXiv.org

This work addresses the challenges of automated cyber defense in complex environments, where large state and action spaces coupled with inefficient exploration hinder traditional deep reinforcement learning approaches due to poor sample efficiency. The problem is formalized as a context-dependent partially observable Markov decision process, and a planning-centric policy is proposed that integrates graph neural networks (GNNs) with Monte Carlo tree search (MCTS). Specifically, the method employs GNNs to generate permutation-invariant embeddings of network topologies and leverages graph-editing action priors to guide MCTS toward more efficient exploration. Policy distillation is further utilized to jointly enable model-free generalization and forward-looking planning. Evaluated on diverse scenarios in CAGE-4, the approach significantly outperforms existing reinforcement learning baselines in both defensive reward and robustness.

1 citationsRead paper
Recent publications

Latest Papers