CORAL: Curriculum-Optimized Reward Adaptation for LiDAR-Based Goal-Directed Urban Driving

๐Ÿ“… 2026-08-14
๐Ÿ“ˆ Citations: 0
โœจ Influential: 0
๐Ÿ“„ PDF
๐Ÿค– AI Summary
This study addresses the challenge ofๅๅŒ learning for multi-objective navigation strategies in urban autonomous driving by proposing the CORAL framework. This approach introduces a novel joint optimization mechanism integrating curriculum learning with dynamic rewards, leveraging polar LiDAR histograms and the PPO algorithm to efficiently learn complex behavioral constraints without relying on point cloud encoders or BEV representations. Experimental results demonstrate that CORAL achieves a 100% success rate on the longest routes and exhibits strong zero-shot transferability to seven new towns with success rates ranging from 68% to 98%, while maintaining lateral deviations below 0.35 m. Consequently, this framework enables efficient training and robust generalization of long-horizon goal-oriented driving policies under complex constraints.
๐Ÿ“ Abstract
Reinforcement learning is promising for autonomous urban driving, but long-horizon goal-directed navigation asks a policy to acquire several competing behaviors at once--reaching a distant goal, tracking a route, avoiding obstacles, obeying signals--and a fixed objective gives no order in which to learn them. This paper presents CORAL, which advances two schedules together: a five-stage curriculum that progressively lengthens routes and tightens behavioral constraints, and a stage-aware reward whose component weights shift emphasis from mission progress toward route following, safety, smoothness, and rule compliance as the task hardens. The policy is a multi-stream actor-critic network trained with Proximal Policy Optimization (PPO) in CARLA on a compact 99-dimensional state pairing a polar LiDAR histogram with vehicle telemetry, ego-frame route geometry, and traffic-rule indicators--no point-cloud encoder, no bird's-eye-view rasterization. Against two PPO baselines under an identical protocol, CORAL reaches the goal in all twenty evaluation episodes on the longest routes under the full set of behavioral constraints, where the baselines reach 5% and 10%; a factorial ablation shows that neither schedule alone matches their combination: removing either lowers both success and route completion, and disabling both drops success to 55%. Trained in one town, the policy transfers zero-shot to seven unseen towns, succeeding in 68-98% of episodes on routes of the same 100-150 m length, with mean lateral deviation below 0.35 m.
Problem

Research questions and friction points this paper is trying to address.

Reinforcement Learning
Autonomous Urban Driving
Goal-Directed Navigation
Reward Adaptation
Curriculum Learning
Innovation

Methods, ideas, or system contributions that make the work stand out.

Curriculum Learning
Reward Adaptation
LiDAR Histogram
Zero-shot Transfer
Reinforcement Learning
๐Ÿ”Ž Similar Papers
2024-04-122024 IEEE Intelligent Vehicles Symposium (IV)Citations: 8
๐Ÿ’ผ Related Jobs
No related jobs found.
A
Anisa Saleem
Department of Computer Engineering, Korea University of Technology and Education (KOREATECH)
Duksu Kim
Duksu Kim
Associate Professor, KOREATECH (Korea University of Technology and Education)
High performance computingCGHDeep learningRobotics