Anomalous Frame Detection by Grouping Frame Similarities between Two Videos Computed by Vision-Language Model to Extract Expert Workers' Unique Actions

📅 2026-07-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of effectively transferring tacit expert skills in critical infrastructure maintenance, which are difficult to capture through conventional methods. The authors propose a novel approach leveraging vision-language models to automatically identify undocumented operational actions by computing frame-level semantic similarity between videos of expert practices and those aligned with standard procedural manuals. By integrating this similarity metric with clustering-based anomaly detection, the method pinpoints deviations indicative of undocumented steps. This work represents the first integration of vision-language models with frame-similarity grouping for skill extraction, overcoming limitations inherent in manual interviews or rule-based systems. Evaluated on switchboard maintenance tasks, the approach successfully uncovered 11 categories of undocumented actions, achieving a knowledge extraction rate of 66.9%—a 50-percentage-point improvement over traditional techniques.
📝 Abstract
Maintenance of critical infrastructures, such as railways and power plants, is essential for operational safety and reliability. However, the declining number of skilled maintenance workers poses a serious challenge to sustaining these operations, highlighting the need to effectively transfer expert know-how to less experienced workers. Although traditional interview-based approaches have been used to elicit maintenance skills, they struggle to capture know-how that experts themselves may not consciously recognize. To address this gap, we proposed a method that detects anomalous frames of candidate actions including know-how by comparing a video of manual-based work with that of expert maintenance workers. In a simulated maintenance experiment involving a distribution board, our method targeted 11 types of actions not described in the manual and achieved a 66.9% extraction rate, marking a 50-percentage-point improvement over conventional techniques. These findings underscore the effectiveness of our approach in revealing hidden maintenance knowledge, thereby contributing to enhanced skill transfer and workforce development in critical infrastructure maintenance.
Problem

Research questions and friction points this paper is trying to address.

anomalous frame detection
expert know-how
skill transfer
vision-language model
maintenance knowledge
Innovation

Methods, ideas, or system contributions that make the work stand out.

anomalous frame detection
vision-language model
expert know-how extraction
skill transfer
maintenance video analysis
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
R
Ryo Sakai
Robotics Research Department, Research and Development Group, Hitachi, Ltd., Ibaraki, Japan
Y
Yongpeng Cao
School of Engineering, The University of Tokyo, Tokyo, Japan
N
Nobutaka Kimura
Next Research Department, Research and Development Group, Hitachi, Ltd., Kokubunji, Japan