Language-Structured Relational Q-Learning for Threat-Aware Control in Safety-Critical Driving

📅 2026-08-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of learning threat-aware and adaptive driving policies in safety-critical scenarios using only observable kinematic information and natural language descriptions. The authors propose a language-structured relational Q-learning approach, introducing language-guided scene modeling into relational reinforcement learning for the first time. They design an ego-centric relational Q-network (ERQ-Net) to jointly learn inter-vehicle dynamic dependencies and action values. The work formally identifies and articulates the “perception-control gap” problem, demonstrating improved performance in 2,500 CARLA simulation scenarios—raising success rates from 49–52% to 55–58% and enhancing adversarial target attention by 1.2–2.1×. Nevertheless, experiments reveal that 76% of scenarios remain solvable by simple policy compositions, highlighting a fundamental performance bottleneck in current methods.
📝 Abstract
Natural-language-based scenario generation offers an intuitive means of describing rare and complex driving interactions, yet it is still uncertain whether training with language-structured data leads to truly adaptive control policies. We propose Language-Structured Relational Q-Learning, instantiated through an Ego-Centric Relational Q-Network (ERQ-Net), which jointly learns inter-vehicle relevance and action values from dynamic traffic graphs. Language descriptions define surrounding-vehicle behaviours during training, while prompts and semantic actor roles are hidden from the policy. ERQ-Net must therefore infer threat relevance solely from observable kinematics and interactions. Across 2,500 safety-critical scenarios, language-structured training improves test success from 49-52% to 55-58% and increases adversary-focused attention from 1.2x to 2.1x, demonstrating emergent threat awareness. However, this representational gain does not consistently translate into adaptive control: trained policies perform similarly to the best constant action, while a portfolio of simple policies solves 76% of scenarios. We formalise this discrepancy as a recognition-control gap and show that reward reweighting and margin shaping do not eliminate the resulting policy collapse. Evaluations of realism, criticality, semantic accuracy, and transfer of state-interface representations to CARLA further highlight both the strengths and the constraints of language-structured relational policy learning in safety-critical driving scenarios.
Problem

Research questions and friction points this paper is trying to address.

threat-aware control
safety-critical driving
language-structured learning
recognition-control gap
adaptive policy
Innovation

Methods, ideas, or system contributions that make the work stand out.

Language-Structured Reinforcement Learning
Relational Q-Learning
Threat-Aware Control
Recognition-Control Gap
Ego-Centric Relational Q-Network
🔎 Similar Papers
💼 Related Jobs
No related jobs found.