CASCADE: A Spatio-Temporal-Causal Reasoning Representation and Dataset for Driving

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文针对自动驾驶中因果推理表示不足的问题,提出CASCADE方法及数据集,通过记录车辆交互的时空因果关系,实现机器可验证的推理预测。
📝 Abstract
Reasoning is a promising route to the generalization that autonomous driving requires in the long tail, as it can infer how the elements of a scene depend on one another and traverse those dependencies to conclusions beyond what is observed. Yet it is hard to tell whether a model's conclusions follow the scene's dependencies, because no driving representation makes them explicit enough to test against. Text-based reasoning traces lack spatio-temporal grounding, spatio-temporal scene graphs lack causal links, and reasoning annotations at scale are increasingly model-generated and hard to verify. To this end, we introduce CASCADE (Causal Spatio-Temporal Analysis of Driving Environments), which encompasses two components: (1) a structured scene representation for reasoning in driving scenes and (2) a human-annotated dataset built on it. For every actor that interacts with the ego vehicle, the CASCADE representation records frame-by-frame, for as long as the actor is visible, what action is taken, where it occurs, and how it depends on the actions and states of others. The resulting structure makes reasoning predictions machine-verifiable: they can be scored against it element by element, without relying on (M)LLM judges. The CASCADE dataset provides comprehensive human annotations for 2,066 driving clips of the PhysicalAI dataset, with over 34K elements that establish the spatio-temporal and causal context of each scene, including 8.6K time-stamped ego and agent actions, 3.7K causal links and 2.9K potential influences, and 6.1K annotations for agents, objects, traffic lights, and environments. Being entirely human-annotated, CASCADE provides the reference for this comparison: benchmarking the reasoning abilities of Physical AI models, and verifying the quality of automatically generated reasoning labels. The CASCADE dataset is available at https://huggingface.co/datasets/nvidia/cascade.
Problem

Research questions and friction points this paper is trying to address.

spatio-temporal
causal reasoning
autonomous driving
scene representation
human-annotated dataset
Innovation

Methods, ideas, or system contributions that make the work stand out.

CASCADE
Spatio-Temporal-Causal Reasoning
Driving Scenes
Human-annotated Dataset
Machine-Verifiable Predictions
🔎 Similar Papers
2024-07-082024 IEEE International Automated Vehicle Validation Conference (IAVVC)Citations: 1