SteelBench: Evaluating Vision-Language Models in Real-World Industrial Environments
Existing video benchmarks struggle to evaluate the capability of vision-language models in recognizing worker behaviors and reasoning about safety rules under real-world industrial surveillance conditions—such as low illumination, occlusion, and long-range viewing. This work proposes SteelBench, the first multidimensional diagnostic benchmark tailored to authentic steel plant environments. Constructed from 149 hours of surveillance footage, it comprises 1,345 densely annotated video clips curated via temporal deduplication, category balancing, and visibility-aware sampling. The dataset encompasses actions, personal protective equipment (PPE) attributes, spatial context, and explicit safety rules, and introduces a novel annotation provenance auditing mechanism. Experiments reveal that even the best-performing model achieves only 42.6% accuracy on action recognition (versus 84.6% for humans), exhibits safety judgment error rates of 37–58%, and fails to pass more than two diagnostic tests. Moreover, unaudited model-generated labels can inflate reported accuracy by up to 17 percentage points for related models.