Institution profile

Honda Research Institute

Industry researchasia · jp
Official website
Research library39linked papers
Opportunities0open roles
Selected work

Representative Papers

SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding

Jul 04, 2026

Existing approaches typically model human actions and gaze in isolation, overlooking their intrinsic coupling in behavior understanding. This work proposes SAGE, a unified framework that, for the first time, jointly models the recognition and future prediction of human-object interaction (HOI) and gaze within an end-to-end trainable architecture applicable to both egocentric and exocentric settings. Built upon Transformers, SAGE integrates gaze cues into spatiotemporal attention mechanisms to enable joint reasoning over current and future action-gaze dynamics. We also introduce Exo-Cook, the first dataset with synchronized HOI and gaze annotations for exocentric videos. Experiments demonstrate that SAGE achieves performance on par with or superior to state-of-the-art methods specialized for individual tasks across VidHOI, EGTEA Gaze+, and Exo-Cook benchmarks.

0 citationsRead paper

What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models

Jun 30, 2026

Autonomous driving systems often struggle to assess the actual impact of occluded traffic participants on motion planning in complex scenarios, leading to either overly conservative behavior or misjudged risks. This work proposes Planning KL Divergence (PKL), a novel metric that leverages vision-language models (VLMs) to rank occluded objects by their planning-criticality and generate structured reasoning annotations. The study introduces the first systematic training framework and benchmark focused on high-impact occlusions, employing a PKL-guided data selection strategy. Experiments on nuScenes demonstrate that a small VLM fine-tuned with PKL-selected data significantly outperforms large zero-shot models, with PKL-based sampling yielding approximately 30% performance improvement over random sampling.

0 citationsRead paper

HOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy Optimization

Jun 15, 2026

This work addresses the challenges of distribution shift, reward misspecification, and per-scenario hyperparameter tuning in robotic motion planning across diverse environments by introducing HOLO-MPPI, a framework that integrates offline high-level policy learning with online low-level stochastic optimal control. The high-level policy, trained via offline reinforcement learning and a world model, generates robust abstract action plans that serve as a conditional sampling prior for Model Predictive Path Integral (MPPI) control. MPPI then performs real-time optimization of low-level controls to handle local disturbances. By embedding a data-driven high-level policy into the MPPI prior, HOLO-MPPI uniquely unifies cross-scenario generalization with real-time adaptability. Experiments demonstrate that HOLO-MPPI significantly outperforms conventional MPPI and end-to-end reinforcement learning baselines across diverse autonomous driving scenarios while maintaining efficient real-time performance.

0 citationsRead paper

A comparative and critical study of EEGNet for fNIRS-driven cognitive load classification

Jun 14, 2026

This study addresses the limited generalization performance in functional near-infrared spectroscopy (fNIRS)-based cognitive workload classification, which stems from temporal variability, inter-subject differences, and sensitivity to preprocessing choices. The authors systematically evaluate the efficacy of the EEGNet architecture for this task by comparing overlapping versus non-overlapping temporal segmentation, window lengths, feature extraction methods (ANOVA, PCA, FastICA), learning rate strategies (fixed vs. adaptive), and evaluation protocols (random split vs. subject-independent). They find that non-overlapping segmentation reduces temporal redundancy and significantly enhances cross-subject generalization. Combining PCA with a 20-second window length yields a subject-independent classification accuracy of 56.11%, establishing a new state-of-the-art result. The work underscores the critical influence of segmentation strategy and learning rate selection on model robustness.

0 citationsRead paper

Driving, Fast or Slow? Neuro-Symbolic Guidance for Motion Prediction in Multi-Modal Ground Mobility

Jun 13, 2026

Existing motion prediction methods often rely on black-box models that struggle to explicitly incorporate traffic rules, resulting in limited interpretability and regulatory compliance. This work proposes the Trajectory Compliance Shaping (TraCS) framework, which innovatively integrates neural networks with symbolic reasoning by translating natural-language traffic rules into probabilistic first-order logic. TraCS dynamically guides prediction models toward compliant trajectories through an agent-driven code generation mechanism and a context-aware confidence decay strategy. Evaluated on the Argoverse 2 benchmark, TraCS consistently enhances the performance of diverse state-of-the-art prediction models, demonstrating the universality, efficiency, and interpretability of probabilistic symbolic reasoning in multimodal ground-vehicle trajectory forecasting.

0 citationsRead paper
Recent publications

Latest Papers

SAGE: Synchronized Action-Gaze Recognition and Anticipation for Human Behavior Understanding

Jul 04, 2026

Existing approaches typically model human actions and gaze in isolation, overlooking their intrinsic coupling in behavior understanding. This work proposes SAGE, a unified framework that, for the first time, jointly models the recognition and future prediction of human-object interaction (HOI) and gaze within an end-to-end trainable architecture applicable to both egocentric and exocentric settings. Built upon Transformers, SAGE integrates gaze cues into spatiotemporal attention mechanisms to enable joint reasoning over current and future action-gaze dynamics. We also introduce Exo-Cook, the first dataset with synchronized HOI and gaze annotations for exocentric videos. Experiments demonstrate that SAGE achieves performance on par with or superior to state-of-the-art methods specialized for individual tasks across VidHOI, EGTEA Gaze+, and Exo-Cook benchmarks.

0 citationsRead paper

What's Hidden Matters: Identifying Planning-Critical Occluded Agents using Vision-Language Models

Jun 30, 2026

Autonomous driving systems often struggle to assess the actual impact of occluded traffic participants on motion planning in complex scenarios, leading to either overly conservative behavior or misjudged risks. This work proposes Planning KL Divergence (PKL), a novel metric that leverages vision-language models (VLMs) to rank occluded objects by their planning-criticality and generate structured reasoning annotations. The study introduces the first systematic training framework and benchmark focused on high-impact occlusions, employing a PKL-guided data selection strategy. Experiments on nuScenes demonstrate that a small VLM fine-tuned with PKL-selected data significantly outperforms large zero-shot models, with PKL-based sampling yielding approximately 30% performance improvement over random sampling.

0 citationsRead paper

HOLO-MPPI: Multi-Scenario Motion Planning via Hierarchical Policy Optimization

Jun 15, 2026

This work addresses the challenges of distribution shift, reward misspecification, and per-scenario hyperparameter tuning in robotic motion planning across diverse environments by introducing HOLO-MPPI, a framework that integrates offline high-level policy learning with online low-level stochastic optimal control. The high-level policy, trained via offline reinforcement learning and a world model, generates robust abstract action plans that serve as a conditional sampling prior for Model Predictive Path Integral (MPPI) control. MPPI then performs real-time optimization of low-level controls to handle local disturbances. By embedding a data-driven high-level policy into the MPPI prior, HOLO-MPPI uniquely unifies cross-scenario generalization with real-time adaptability. Experiments demonstrate that HOLO-MPPI significantly outperforms conventional MPPI and end-to-end reinforcement learning baselines across diverse autonomous driving scenarios while maintaining efficient real-time performance.

0 citationsRead paper

A comparative and critical study of EEGNet for fNIRS-driven cognitive load classification

Jun 14, 2026

This study addresses the limited generalization performance in functional near-infrared spectroscopy (fNIRS)-based cognitive workload classification, which stems from temporal variability, inter-subject differences, and sensitivity to preprocessing choices. The authors systematically evaluate the efficacy of the EEGNet architecture for this task by comparing overlapping versus non-overlapping temporal segmentation, window lengths, feature extraction methods (ANOVA, PCA, FastICA), learning rate strategies (fixed vs. adaptive), and evaluation protocols (random split vs. subject-independent). They find that non-overlapping segmentation reduces temporal redundancy and significantly enhances cross-subject generalization. Combining PCA with a 20-second window length yields a subject-independent classification accuracy of 56.11%, establishing a new state-of-the-art result. The work underscores the critical influence of segmentation strategy and learning rate selection on model robustness.

0 citationsRead paper

Driving, Fast or Slow? Neuro-Symbolic Guidance for Motion Prediction in Multi-Modal Ground Mobility

Jun 13, 2026

Existing motion prediction methods often rely on black-box models that struggle to explicitly incorporate traffic rules, resulting in limited interpretability and regulatory compliance. This work proposes the Trajectory Compliance Shaping (TraCS) framework, which innovatively integrates neural networks with symbolic reasoning by translating natural-language traffic rules into probabilistic first-order logic. TraCS dynamically guides prediction models toward compliant trajectories through an agent-driven code generation mechanism and a context-aware confidence decay strategy. Evaluated on the Argoverse 2 benchmark, TraCS consistently enhances the performance of diverse state-of-the-art prediction models, demonstrating the universality, efficiency, and interpretability of probabilistic symbolic reasoning in multimodal ground-vehicle trajectory forecasting.

0 citationsRead paper