Anomalous Frame Detection Using VLM-Based Description Comparison for Extracting Expert-Specific Actions and Contextual Decision-Making Scenes with Intra-Video Self-Similarity

📅 2026-07-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the shortage of skilled maintenance personnel by proposing a novel method to automatically extract critical scenes from expert demonstration and standard operating procedure videos, thereby capturing experts’ distinctive action patterns and implicit contextual decision-making knowledge. The approach uniquely integrates vision-language models to generate frame-level descriptions and combines cross-video description contrast with intra-video self-similarity analysis to identify divergent operational and decision-making scenarios through anomaly frame detection. Evaluated on 27 switchboard maintenance tasks, the method achieves extraction rates of 65% for action-related scenes and 61% for decision-related scenes—significantly outperforming conventional approaches, which attain only 59% and 33%, respectively—demonstrating its effectiveness and innovation.
📝 Abstract
Maintenance of critical infrastructures, such as railways and power plants, is essential for ensuring operational safety and reliability. However, the declining number of skilled maintenance workers highlights the need to transfer expert know-how to less experienced workers. Previous studies have attempted to extract candidates of expert knowledge by comparing videos of manual-based work with those of expert workers, mainly focusing on differences in observable actions. However, expert know-how is often embedded not only in actions but also in contextual decision-making during task execution. This paper proposes a method that detects anomalous frames between two task videos to automatically extract candidate scenes containing expert-specific actions and contextual decision-making scenes. The method generates frame-wise visual descriptions using a vision-language model (VLM). Expert-specific actions are extracted based on frame similarities computed from description comparisons between two videos, while contextual decision-making scenes are extracted using segment similarities derived from intra-video self-similarity of the descriptions. In simulated distribution board maintenance experiments involving 27 task scenarios, the proposed method achieved extraction rates of 65% for action candidates and 61% for decision-scene candidates, improving over conventional methods that achieved 59% and 33%, respectively. These results demonstrate the effectiveness of the proposed approach in discovering candidate scenes containing expert know-how.
Problem

Research questions and friction points this paper is trying to address.

expert knowledge transfer
anomalous frame detection
contextual decision-making
vision-language model
intra-video self-similarity
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
anomalous frame detection
expert knowledge extraction
intra-video self-similarity
contextual decision-making
💼 Related Jobs
No related jobs found.
R
Ryo Sakai
Robotics Research Department, Research and Development Group, Hitachi, Ltd., Ibaraki, Japan
K
Kaname Yokoyama
Robotics Research Department, Research and Development Group, Hitachi, Ltd., Ibaraki, Japan