LLM-Based Hierarchical Coordinated Control with Continuation-Aware Policy Learning

📅 2026-08-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenges of multi-unit collaborative modeling, information heterogeneity, and strict action constraints in complex engineering systems by proposing a large language model-based hierarchical collaboration framework. The core innovation lies in the design of a Continuity-Aware GRPO algorithm, which evaluates the evolutionary impact of decisions on subsequent control intervals to effectively resolve heterogeneous context reasoning and constraint enforcement difficulties. By integrating large language models with reinforcement learning, this approach achieves precise cross-level decision-making. Empirical validation in traffic flow control and virtual power plant tasks demonstrates that the proposed method comprehensively outperforms existing baselines, significantly enhancing collaborative control efficacy in complex systems.
📝 Abstract
Coordinating multiple interacting units in complex engineering systems is challenging when system interactions are difficult to model, operational information is heterogeneous, and low-level actions must satisfy strict constraints. We propose an LLM-based hierarchical framework in which the LLM coordinates interacting units based on heterogeneous operational context, while task-specific controllers or optimizers generate executable and constraint-aware actions. We further introduce Continuation-Aware GRPO to capture the consequences of coordination decisions over subsequent control intervals. Rather than judging a decision only by its immediate outcome, the method also evaluates how the system evolves afterward under the current policy. We validate the framework on multi-ramp traffic control and virtual power plant (VPP) energy management, using simplified system models for training and more realistic simulators for evaluation. Across both tasks, the proposed method consistently outperforms direct task-specific control and optimization, end-to-end reinforcement learning, rule-based and RL-based hierarchical coordination, and prompting-only LLM coordinators, demonstrating the value of heterogeneous-context reasoning, hierarchical execution, and continuation-aware policy learning.
Problem

Research questions and friction points this paper is trying to address.

Multi-agent Coordination
Complex Engineering Systems
Heterogeneous Information
Constraint Satisfaction
Hierarchical Control
Innovation

Methods, ideas, or system contributions that make the work stand out.

Hierarchical Coordinated Control
Continuation-Aware GRPO
LLM-Based Coordination
Heterogeneous Context Reasoning
Constraint-Aware Actions
🔎 Similar Papers
No similar papers found.
C
Changhong He
Beihang University, Beijing, China; Baidu, Inc., Beijing, China
J
Jinda Gao
Baidu, Inc., Beijing, China
X
Xinkuan Liu
Baidu, Inc., Beijing, China
L
Le Zhang
Baidu, Inc., Beijing, China
X
Xizi Luo
Beihang University, Beijing, China; Baidu, Inc., Beijing, China
Yu Mei
Yu Mei
Michigan State University
Soft RoboticsControl