Institution profile

Nexar

Industry researchnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Beyond the Beep: Scalable Collision Anticipation and Real-Time Explainability with BADAS-2.0

Apr 07, 2026

This work addresses the limited capability of advanced driver-assistance systems to anticipate collisions in long-tail, high-risk driving scenarios by proposing an efficient and interpretable real-time prediction framework. Leveraging the Nexar Atlas platform, the authors curate a large-scale dataset comprising 178,500 annotated video clips. The approach integrates V-JEPA2 self-supervised pretraining, active learning for targeted annotation of hazardous scenarios, edge-device-oriented knowledge distillation, and a vision-language model (BADAS-Reason) that generates object-level attention heatmaps alongside natural language rationales. The resulting models are compressed to 22–86 million parameters, achieving 7–12× inference speedup while preserving accuracy, thereby significantly enhancing both predictive performance and interpretability in long-tail safety-critical situations.

0 citationsRead paper

Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

Mar 09, 2025

This work addresses the insufficient interpretability of vision-language models (VLMs) in high-stakes applications. We propose PixelSHAP—the first model-agnostic, pixel-level attribution framework specifically designed for VLMs. PixelSHAP extends Shapley value theory to structured visual entities by introducing zero-shot object masking and recomposition, embedding-space similarity measurement, and an efficient sampling strategy to quantify each pixel’s contribution to VLM outputs—thereby jointly supporting semantic object-level understanding and precise pixel-level localization. Crucially, it operates without access to model parameters or gradients, enabling black-box interpretation of commercial VLMs. Evaluated on real-world high-risk scenarios—including autonomous driving—PixelSHAP successfully identifies critical visual drivers underlying model decisions, substantially enhancing transparency and trust. An open-source implementation demonstrates strong cross-model robustness and practical utility across diverse VLM architectures.

0 citationsRead paper

Nexar Dashcam Collision Prediction Dataset and Challenge

Mar 05, 2025

This work addresses the early prediction of imminent collisions in real-world traffic scenarios. We introduce the first fine-grained collision prediction dataset comprising 1,500 real 40-second videos, annotated with collision event types, environmental and scene attributes, precise event timestamps, and—novelly—the “predictable moment,” i.e., the earliest time at which a model can reliably issue a warning. To holistically evaluate timeliness and reliability, we propose a joint average precision (AP) metric across multiple lead times (500 ms, 1,000 ms, and 1,500 ms). Our method integrates video understanding, temporal action detection, and multi-task event forecasting. The dataset is publicly released under ethical usage constraints. Benchmark models achieve significant AP improvements at the 1,500-ms lead time, advancing research on high-temporal-precision, deployable safety warning systems for autonomous driving.

0 citationsRead paper
Recent publications

Latest Papers

Beyond the Beep: Scalable Collision Anticipation and Real-Time Explainability with BADAS-2.0

Apr 07, 2026

This work addresses the limited capability of advanced driver-assistance systems to anticipate collisions in long-tail, high-risk driving scenarios by proposing an efficient and interpretable real-time prediction framework. Leveraging the Nexar Atlas platform, the authors curate a large-scale dataset comprising 178,500 annotated video clips. The approach integrates V-JEPA2 self-supervised pretraining, active learning for targeted annotation of hazardous scenarios, edge-device-oriented knowledge distillation, and a vision-language model (BADAS-Reason) that generates object-level attention heatmaps alongside natural language rationales. The resulting models are compressed to 22–86 million parameters, achieving 7–12× inference speedup while preserving accuracy, thereby significantly enhancing both predictive performance and interpretability in long-tail safety-critical situations.

0 citationsRead paper

Attention, Please! PixelSHAP Reveals What Vision-Language Models Actually Focus On

Mar 09, 2025

This work addresses the insufficient interpretability of vision-language models (VLMs) in high-stakes applications. We propose PixelSHAP—the first model-agnostic, pixel-level attribution framework specifically designed for VLMs. PixelSHAP extends Shapley value theory to structured visual entities by introducing zero-shot object masking and recomposition, embedding-space similarity measurement, and an efficient sampling strategy to quantify each pixel’s contribution to VLM outputs—thereby jointly supporting semantic object-level understanding and precise pixel-level localization. Crucially, it operates without access to model parameters or gradients, enabling black-box interpretation of commercial VLMs. Evaluated on real-world high-risk scenarios—including autonomous driving—PixelSHAP successfully identifies critical visual drivers underlying model decisions, substantially enhancing transparency and trust. An open-source implementation demonstrates strong cross-model robustness and practical utility across diverse VLM architectures.

0 citationsRead paper

Nexar Dashcam Collision Prediction Dataset and Challenge

Mar 05, 2025

This work addresses the early prediction of imminent collisions in real-world traffic scenarios. We introduce the first fine-grained collision prediction dataset comprising 1,500 real 40-second videos, annotated with collision event types, environmental and scene attributes, precise event timestamps, and—novelly—the “predictable moment,” i.e., the earliest time at which a model can reliably issue a warning. To holistically evaluate timeliness and reliability, we propose a joint average precision (AP) metric across multiple lead times (500 ms, 1,000 ms, and 1,500 ms). Our method integrates video understanding, temporal action detection, and multi-task event forecasting. The dataset is publicly released under ethical usage constraints. Benchmark models achieve significant AP improvements at the 1,500-ms lead time, advancing research on high-temporal-precision, deployable safety warning systems for autonomous driving.

0 citationsRead paper