Institution profile

Centre for Artificial Intelligence Research

Academic institution
Official website
Research library2linked papers
Opportunities0open roles
Selected work

Representative Papers

Limitations of Scalarisation in MORL: A Comparative Study in Discrete Environments

Nov 20, 2025

This study systematically exposes fundamental limitations of scalarization-based methods in multi-objective reinforcement learning (MORL) with discrete action and observation spaces: poor Pareto-front coverage, low robustness, and strong dependence on environmental properties and front geometry. To address these issues, we propose an inner-loop multi-policy architecture and comparatively evaluate three representative approaches—linear scalarization, Chebyshev scalarization, and non-scalarized Pareto Q-learning—under both outer-loop single-policy and inner-loop multi-policy paradigms. Results demonstrate that Pareto Q-learning significantly improves solution-set diversity and stability over scalarized methods. Moreover, the inner-loop multi-policy design effectively mitigates scalarization’s sensitivity to weight selection and susceptibility to local optima. Our empirical analysis provides both a novel methodological paradigm and rigorous evidence for developing robust, scalable MORL algorithms.

0 citationsRead paper

A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms

Nov 20, 2025

This study investigates how decision transformers (DTs) compare to conventional offline reinforcement learning (RL) algorithms—specifically conservative Q-learning (CQL) and implicit Q-learning (IQL)—under varying reward densities (dense vs. sparse) in the ANT continuous-control benchmark. Method: We conduct a systematic, controlled empirical evaluation across uniformly configured offline datasets of varying quality and reward sparsity. Contribution/Results: We find that DTs exhibit remarkable robustness to reward density shifts: they outperform both CQL and IQL in sparse-reward regimes and under medium-quality offline data, achieving higher policy performance, greater stability, and lower evaluation variance. In contrast, IQL excels in dense-reward settings, while CQL demonstrates superior overall robustness across diverse conditions. Crucially, this work provides the first empirical evidence that sequence modeling—via autoregressive action prediction—confers distinct advantages in low signal-to-noise-ratio feedback environments. These findings offer principled guidance for reward-structure-aware algorithm selection and design in offline RL.

0 citationsRead paper
Recent publications

Latest Papers

Limitations of Scalarisation in MORL: A Comparative Study in Discrete Environments

Nov 20, 2025

This study systematically exposes fundamental limitations of scalarization-based methods in multi-objective reinforcement learning (MORL) with discrete action and observation spaces: poor Pareto-front coverage, low robustness, and strong dependence on environmental properties and front geometry. To address these issues, we propose an inner-loop multi-policy architecture and comparatively evaluate three representative approaches—linear scalarization, Chebyshev scalarization, and non-scalarized Pareto Q-learning—under both outer-loop single-policy and inner-loop multi-policy paradigms. Results demonstrate that Pareto Q-learning significantly improves solution-set diversity and stability over scalarized methods. Moreover, the inner-loop multi-policy design effectively mitigates scalarization’s sensitivity to weight selection and susceptibility to local optima. Our empirical analysis provides both a novel methodological paradigm and rigorous evidence for developing robust, scalable MORL algorithms.

0 citationsRead paper

A Comparison Between Decision Transformers and Traditional Offline Reinforcement Learning Algorithms

Nov 20, 2025

This study investigates how decision transformers (DTs) compare to conventional offline reinforcement learning (RL) algorithms—specifically conservative Q-learning (CQL) and implicit Q-learning (IQL)—under varying reward densities (dense vs. sparse) in the ANT continuous-control benchmark. Method: We conduct a systematic, controlled empirical evaluation across uniformly configured offline datasets of varying quality and reward sparsity. Contribution/Results: We find that DTs exhibit remarkable robustness to reward density shifts: they outperform both CQL and IQL in sparse-reward regimes and under medium-quality offline data, achieving higher policy performance, greater stability, and lower evaluation variance. In contrast, IQL excels in dense-reward settings, while CQL demonstrates superior overall robustness across diverse conditions. Crucially, this work provides the first empirical evidence that sequence modeling—via autoregressive action prediction—confers distinct advantages in low signal-to-noise-ratio feedback environments. These findings offer principled guidance for reward-structure-aware algorithm selection and design in offline RL.

0 citationsRead paper