Institution profile

Hitachi, Ltd.

Industry researchasia · jp
Official website
Research library47linked papers
Opportunities0open roles
Selected work

Representative Papers

AI Evaluation Should Measure Verification Cost, Not Correctness Alone

Aug 09, 2026

Current AI evaluation frameworks overemphasize output correctness while neglecting the resource costs required to verify errors in real-world deployment, allowing high accuracy metrics to mask substantial verification burdens. This work introduces verification-cost errors (VCEs)—errors that cannot be detected by a specified proportion of validators within a given verification budget—thereby shifting the paradigm from defining errors solely by output properties to centering on their detectability during verification. Through an operational definition, verification budget modeling, and user studies, we empirically demonstrate in code generation and multimodal document understanding tasks that high benchmark accuracy can coexist with significant verification effort, underscoring that correctness alone is insufficient to reflect system reliability in practical settings.

0 citationsRead paper

AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching

Aug 03, 2026

Existing motion skill coaching methods rely on real-time expert demonstrations, resulting in high deployment costs and limited scalability. This work proposes AIDE, a novel framework that, for the first time, leverages expert demonstrations only during training via knowledge distillation and, at inference time, generates high-quality natural language feedback using solely the learner’s pose sequence—eliminating any dependence on real-time expert data. AIDE employs a teacher–student architecture: the teacher model, built upon a frozen large language model, extracts discrepancy encodings from paired learner–expert poses; the student model inherits this encoder and incorporates an auxiliary module to predict complementary information from a single-stream input. Evaluated on the ExpertAF dataset, AIDE significantly outperforms reference-free baselines and matches the performance of two-stage methods requiring expert demonstrations, with further validation provided by LLM-based assessments.

0 citationsRead paper

MAGE-Vein: Multi-Instance Age and Gender Estimation from Finger Vein Images

Jul 22, 2026

Finger vein imaging has long been considered unsuitable for age estimation, primarily due to demographic biases in publicly available datasets and confounding physiological factors such as gender. This work proposes a multi-instance, multi-task learning framework that performs feature-level fusion across three fingers to extract structured aging-related features, while jointly optimizing gender classification to mitigate gender-specific vascular variations. For the first time, the study demonstrates high-accuracy age estimation from finger vein images on a balanced dataset of 402 subjects, achieving a mean absolute error of 6.12 years and a correlation coefficient of 0.880. These results challenge the prevailing consensus regarding the modality’s inadequacy, revealing that prior failures stemmed from data bias rather than inherent limitations of finger vein imaging itself.

0 citationsRead paper

Anomalous Frame Detection by Grouping Frame Similarities between Two Videos Computed by Vision-Language Model to Extract Expert Workers' Unique Actions

Jul 12, 2026

This study addresses the challenge of effectively transferring tacit expert skills in critical infrastructure maintenance, which are difficult to capture through conventional methods. The authors propose a novel approach leveraging vision-language models to automatically identify undocumented operational actions by computing frame-level semantic similarity between videos of expert practices and those aligned with standard procedural manuals. By integrating this similarity metric with clustering-based anomaly detection, the method pinpoints deviations indicative of undocumented steps. This work represents the first integration of vision-language models with frame-similarity grouping for skill extraction, overcoming limitations inherent in manual interviews or rule-based systems. Evaluated on switchboard maintenance tasks, the approach successfully uncovered 11 categories of undocumented actions, achieving a knowledge extraction rate of 66.9%—a 50-percentage-point improvement over traditional techniques.

0 citationsRead paper

Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

Jul 12, 2026

This study addresses the shortage of skilled maintenance personnel by proposing a novel method to automatically extract critical scenes from expert demonstration and standard operating procedure videos, thereby capturing experts’ distinctive action patterns and implicit contextual decision-making knowledge. The approach uniquely integrates vision-language models to generate frame-level descriptions and combines cross-video description contrast with intra-video self-similarity analysis to identify divergent operational and decision-making scenarios through anomaly frame detection. Evaluated on 27 switchboard maintenance tasks, the method achieves extraction rates of 65% for action-related scenes and 61% for decision-related scenes—significantly outperforming conventional approaches, which attain only 59% and 33%, respectively—demonstrating its effectiveness and innovation.

0 citationsRead paper
Recent publications

Latest Papers

AI Evaluation Should Measure Verification Cost, Not Correctness Alone

Aug 09, 2026

Current AI evaluation frameworks overemphasize output correctness while neglecting the resource costs required to verify errors in real-world deployment, allowing high accuracy metrics to mask substantial verification burdens. This work introduces verification-cost errors (VCEs)—errors that cannot be detected by a specified proportion of validators within a given verification budget—thereby shifting the paradigm from defining errors solely by output properties to centering on their detectability during verification. Through an operational definition, verification budget modeling, and user studies, we empirically demonstrate in code generation and multimodal document understanding tasks that high benchmark accuracy can coexist with significant verification effort, underscoring that correctness alone is insufficient to reflect system reliability in practical settings.

0 citationsRead paper

AIDE: Automated Instruction via Distilled Expertise for Reference-Free Motor Skill Coaching

Aug 03, 2026

Existing motion skill coaching methods rely on real-time expert demonstrations, resulting in high deployment costs and limited scalability. This work proposes AIDE, a novel framework that, for the first time, leverages expert demonstrations only during training via knowledge distillation and, at inference time, generates high-quality natural language feedback using solely the learner’s pose sequence—eliminating any dependence on real-time expert data. AIDE employs a teacher–student architecture: the teacher model, built upon a frozen large language model, extracts discrepancy encodings from paired learner–expert poses; the student model inherits this encoder and incorporates an auxiliary module to predict complementary information from a single-stream input. Evaluated on the ExpertAF dataset, AIDE significantly outperforms reference-free baselines and matches the performance of two-stage methods requiring expert demonstrations, with further validation provided by LLM-based assessments.

0 citationsRead paper

MAGE-Vein: Multi-Instance Age and Gender Estimation from Finger Vein Images

Jul 22, 2026

Finger vein imaging has long been considered unsuitable for age estimation, primarily due to demographic biases in publicly available datasets and confounding physiological factors such as gender. This work proposes a multi-instance, multi-task learning framework that performs feature-level fusion across three fingers to extract structured aging-related features, while jointly optimizing gender classification to mitigate gender-specific vascular variations. For the first time, the study demonstrates high-accuracy age estimation from finger vein images on a balanced dataset of 402 subjects, achieving a mean absolute error of 6.12 years and a correlation coefficient of 0.880. These results challenge the prevailing consensus regarding the modality’s inadequacy, revealing that prior failures stemmed from data bias rather than inherent limitations of finger vein imaging itself.

0 citationsRead paper

Anomalous Frame Detection by Grouping Frame Similarities between Two Videos Computed by Vision-Language Model to Extract Expert Workers' Unique Actions

Jul 12, 2026

This study addresses the challenge of effectively transferring tacit expert skills in critical infrastructure maintenance, which are difficult to capture through conventional methods. The authors propose a novel approach leveraging vision-language models to automatically identify undocumented operational actions by computing frame-level semantic similarity between videos of expert practices and those aligned with standard procedural manuals. By integrating this similarity metric with clustering-based anomaly detection, the method pinpoints deviations indicative of undocumented steps. This work represents the first integration of vision-language models with frame-similarity grouping for skill extraction, overcoming limitations inherent in manual interviews or rule-based systems. Evaluated on switchboard maintenance tasks, the approach successfully uncovered 11 categories of undocumented actions, achieving a knowledge extraction rate of 66.9%—a 50-percentage-point improvement over traditional techniques.

0 citationsRead paper

Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

Jul 12, 2026

This study addresses the shortage of skilled maintenance personnel by proposing a novel method to automatically extract critical scenes from expert demonstration and standard operating procedure videos, thereby capturing experts’ distinctive action patterns and implicit contextual decision-making knowledge. The approach uniquely integrates vision-language models to generate frame-level descriptions and combines cross-video description contrast with intra-video self-similarity analysis to identify divergent operational and decision-making scenarios through anomaly frame detection. Evaluated on 27 switchboard maintenance tasks, the method achieves extraction rates of 65% for action-related scenes and 61% for decision-related scenes—significantly outperforming conventional approaches, which attain only 59% and 33%, respectively—demonstrating its effectiveness and innovation.

0 citationsRead paper