Institution profile

University College London

Academic institutioneurope · gb
Official website
Research library2,218linked papers
Opportunities0open roles
Selected work

Representative Papers

Identification and Estimation of Network Models with Nonparametric Unobserved Heterogeneity

Feb 06, 2026

Unobserved heterogeneity in network models—such as latent homophily—often leads to biased estimation and misleading policy implications. This study proposes a nonparametric approach that identifies individuals sharing the same unobserved fixed effects through their interactive outcomes, leveraging variation in their observed characteristics to consistently estimate covariate effects without imposing parametric assumptions on the functional form of the fixed effects. Relying solely on interaction data, the method integrates nonparametric identification, fixed-effect control, and large-sample asymptotic theory to yield an estimator with strong theoretical properties. Numerical simulations confirm the estimator’s effectiveness and robustness in handling complex unobserved heterogeneity.

17 citations2 influentialRead paper

Enhancing efficiency and propulsion in bio-mimetic robotic fish through end-to-end deep reinforcement learning

Mar 01, 2024The Physics of Fluids

Bionic robotic fish suffer from low propulsion efficiency and high energy consumption. Method: This study proposes an end-to-end deep reinforcement learning (DRL) control framework, introducing— for the first time in underwater bionic robotics—extended pressure sensing combined with temporal Transformer modeling, integrated with a policy transfer mechanism to enhance training stability and environmental adaptability. Training achieves autonomous, stable, and rapid convergence within CFD simulations (Re = 6000). Contribution/Results: The DRL policy improves propulsion efficiency by 37% and reduces specific energy consumption per unit thrust by 29% over conventional pre-programmed gaits. Flow-field analysis reveals that efficiency stems from embodied regulation of body deformation and vortex–body interactions. The core contribution is a novel bio-inspired locomotion control paradigm unifying perception, spatiotemporal modeling, and decision-making.

9 citationsRead paper

Right Now, Wrong Then: Non-Stationary Direct Preference Optimization under Preference Drift

Jul 26, 2024arXiv.org

Existing RLHF methods assume static human preferences, neglecting their non-stationary drift induced by societal changes and environmental evolution—leading to alignment degradation. This paper proposes Non-Stationary Direct Preference Optimization (NS-DPO), the first framework to model preference dynamics via a time-varying Bradley–Terry model and incorporate an exponentially time-weighted DPO loss. We theoretically derive an upper bound on estimation error under non-stationarity and prove offline convergence. NS-DPO integrates time-aware preference modeling, exponential discounting, and rigorous theoretical analysis. Experiments on simulated preference-drift scenarios demonstrate that NS-DPO significantly outperforms standard DPO and other baselines, achieving both robustness to drift and zero performance degradation in stationary settings. It provides the first solution for aligning LLMs under time-varying preferences with provable theoretical guarantees and empirical effectiveness.

4 citationsRead paper

Policy choice in time series by empirical welfare maximization

May 08, 2022

Dynamic multivariate time series pose challenges including time-varying environments, historical dependence, dynamic causal effects, and statistical dependencies. Method: This paper proposes Time-series Empirical Welfare Maximization (T-EWM), the first extension of the empirical welfare maximization framework to dynamic time-series settings. T-EWM employs nonparametric potential outcome modeling and conditional welfare optimization to learn dynamic optimal policies. Contribution/Results: We establish theoretical guarantees, including conditional welfare consistency and a non-asymptotic upper bound on policy regret. In simulation studies and an empirical application to COVID-19 containment policy evaluation, T-EWM significantly improves policy welfare and achieves rapid regret convergence under limited samples. The framework provides a novel paradigm for dynamic decision-making that balances statistical rigor with practical feasibility.

4 citationsRead paper

Improving Equivariant Networks with Probabilistic Symmetry Breaking

Mar 27, 2025

Equivariant networks strictly preserve input symmetries, rendering them ill-suited for generative tasks requiring *active symmetry breaking*—e.g., reconstructing asymmetric structures from highly symmetric latent representations. To address this, we establish the first necessary and sufficient representation theorem for equivariant conditional distributions and propose SymPE: a method that achieves *controllable symmetry breaking* via learnable stochastic normalized positional encodings, while preserving the group-equivariant inductive bias. SymPE unifies probabilistic symmetry breaking, positional encoding, and equivariant graph neural networks, and naturally integrates with diffusion-based generative frameworks. Empirically, it significantly improves performance on graph diffusion modeling, graph autoencoding, and lattice spin system generation. Theoretically, we prove that SymPE’s generalization bound is strictly superior to that of conventional equivariant networks.

3 citationsRead paper
Recent publications

Latest Papers