Vision-Language Models for Occupational Physical Exposure Assessment: Estimating External Hand Forces in Manual Material Handling Tasks from RGB Video

📅 2026-08-23
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究使用视觉-语言模型从RGB视频估计手动物料搬运任务中的外部手部力,以解决职业物理暴露评估中连续力测量的问题。
📝 Abstract
External hand forces are important inputs to biomechanical analyses of occupational physical exposure and injury risk, yet continuous force measurements during manual material handling (MMH) typically requires instrumented objects or specialized sensing. We evaluated a vision-language model (VLM)-based pipeline that combines task-specific textual cues, visual representations, and known box mass to estimate dynamic, triaxial, bilateral external hand forces from RGB video. Thirty-five healthy young adults performed five MMH tasks involving lifting, carrying, pushing, and pulling with box masses of 6, 9, and 12 kg. The pipeline used text-guided localization of participant and handled-object regions of interest (ROIs), pretrained vision-transformer feature extraction, and transformer-based temporal regression. Performance was evaluated using leave-one-subject-out validation across seven camera-view conditions (three single-view and four multi-view conditions) and four ROI strategies. Overall, root mean square error was ~4.7-5.6 N for the horizontal and mediolateral force components and ~10.6-11.0 N for the vertical component. Including the handled object as a second ROI generally improved force estimation, with some of the largest benefits under single-camera conditions, whereas pixel-level segmentation provided little additional improvement. Multi-camera capture provided the clearest benefit for peak-force estimation, particularly for the vertical component, whereas differences in overall frame-level error among camera configurations were comparatively modest. These findings demonstrate the feasibility of estimating continuous, bilateral, directional hand-force estimates from RGB video and known load mass without requiring sensors on the worker or handled objects as model inputs, supporting the development of more scalable occupational physical exposure and risk assessments.
Problem

Research questions and friction points this paper is trying to address.

external hand forces
manual material handling
RGB video
vision-language model
occupational physical exposure
Innovation

Methods, ideas, or system contributions that make the work stand out.

vision-language model
external hand forces
RGB video
occupational physical exposure
task-specific textual cues
💼 Related Jobs
No related jobs found.
M
Mohammad Sadra Rajabi
Department of Industrial and Systems Engineering, Virginia Tech, Blacksburg VA 24061, USA
A
Aanuoluwapo Ojelade
St. Jude Children's Research Hospital, Memphis, TN 38105, USA
S
Sunwook Kim
Department of Industrial and Systems Engineering, Virginia Tech, Blacksburg VA 24061, USA
M
Maury A. Nussbaum
Department of Industrial and Systems Engineering, Virginia Tech, Blacksburg VA 24061, USA