Enhancing Human-Likeness in Reinforcement Learning Agents via Hierarchical Macro Action Quantization
Reinforcement learning agents often lack interpretability and reliability due to behavioral discrepancies from humans. To address this, this work proposes Hierarchical Macro-Action Quantization (HiMAQ), a method that encodes human demonstrations into structured behavioral units through a two-level vector quantization mechanism: first clustering fine-grained sub-actions and then aggregating them into high-level macro-actions. Experiments on the D4RL benchmark demonstrate that HiMAQ consistently enhances the human-likeness of agent behavior across multiple offline reinforcement learning algorithms—including IQL, SAC, and RLPD—while maintaining comparable or higher task success rates. HiMAQ outperforms its non-hierarchical counterpart, MAQ, exhibiting strong generalization capability and practical utility.