Robusto-1 Dataset: Comparing Humans and VLMs on real out-of-distribution Autonomous Driving VQA from Peru

📅 2025-03-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study investigates the cognitive response alignment between foundational vision-language models (VLMs) and human drivers under out-of-distribution (OoD) autonomous driving scenarios. Method: We introduce the first cross-distribution visual question answering (VQA) benchmark grounded in real-world Peruvian traffic conditions—featuring non-synthetic, high-difficulty anomalous traffic videos—and pioneer the application of representational similarity analysis (RSA), a neuroscience-inspired technique, to quantify fine-grained human-AI cognitive alignment in multimodal VQA evaluation. Results: Human-VLM response consistency is strongly problem-type-dependent; systematic misalignment emerges in high-level cognitive tasks such as anomalous object recognition and driver intent inference, exposing fundamental limitations of current VLMs in realistic, long-tailed driving environments. Our core contributions are (1) a novel, interpretable human-AI cognitive comparison paradigm, and (2) the first OoD VQA benchmark explicitly targeting complex, infrastructure-constrained road conditions in developing countries.

Technology Category

Application Category

📝 Abstract
As multimodal foundational models start being deployed experimentally in Self-Driving cars, a reasonable question we ask ourselves is how similar to humans do these systems respond in certain driving situations -- especially those that are out-of-distribution? To study this, we create the Robusto-1 dataset that uses dashcam video data from Peru, a country with one of the worst (aggressive) drivers in the world, a high traffic index, and a high ratio of bizarre to non-bizarre street objects likely never seen in training. In particular, to preliminarly test at a cognitive level how well Foundational Visual Language Models (VLMs) compare to Humans in Driving, we move away from bounding boxes, segmentation maps, occupancy maps or trajectory estimation to multi-modal Visual Question Answering (VQA) comparing both humans and machines through a popular method in systems neuroscience known as Representational Similarity Analysis (RSA). Depending on the type of questions we ask and the answers these systems give, we will show in what cases do VLMs and Humans converge or diverge allowing us to probe on their cognitive alignment. We find that the degree of alignment varies significantly depending on the type of questions asked to each type of system (Humans vs VLMs), highlighting a gap in their alignment.
Problem

Research questions and friction points this paper is trying to address.

Compares human and VLM responses in autonomous driving scenarios.
Evaluates cognitive alignment using Visual Question Answering (VQA).
Highlights gaps in understanding out-of-distribution driving situations.
Innovation

Methods, ideas, or system contributions that make the work stand out.

Uses dashcam video data from Peru
Employs Visual Question Answering (VQA)
Applies Representational Similarity Analysis (RSA)
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
D
Dunant Cusipuma
Artificio
D
David Ortega
Artificio
V
Victor Flores-Benites
Artificio, Universidad de Ingeneria y Tecnologia (UTEC)
A
Arturo Deza
Artificio, Universidad de Ingeneria y Tecnologia (UTEC)