π€ AI Summary
This study addresses the limitations of manual, observer-based assessment of patient actions in cognitive rehabilitation for schizophrenia, which suffers from high subjectivity and poor scalability. To overcome these challenges, this work proposes the first automated evaluation platform leveraging a fine-tuned vision-language model. Operating within a customized tabletop miniature environment, the system integrates video recordings with audio instructions to enable end-to-end mapping from handβobject motion trajectories to high-level clinical semantics. By analyzing video sequences, tracking movements, and generating semantic interpretations, the framework automatically validates the correctness of goal-directed behaviors. Evaluated on a dataset comprising 4,634 videos, the method achieves high-precision semantic-level assessment, substantially enhancing objectivity and scalability, and offers an innovative automated solution for cognitive rehabilitation.
π Abstract
Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning. For patients diagnosed with schizophrenia, these tasks are crucial for addressing severe cognitive deficits. However, evaluating the correctness of these physical actions generally relies on manual observation by clinicians, which introduces subjectivity and limits the scalability of therapeutic interventions. In this paper, we propose an automated framework based on Vision-Language Models for action verification in cognitive remediation tasks tailored for schizophrenia rehabilitation. The proposed system relies on a camera-monitored tabletop environment composed of structured miniature scenes including roads, a roundabout, a park, and toy vehicles. Patients receive audio instructions describing goal-oriented spatial actions to perform by manipulating a toy vehicle. These interactive physical activities are specifically designed to stimulate targeted cognitive functions, such as sustained attention, motor coordination, spatial navigation, and cognitive flexibility. To verify the correctness of the performed actions without requiring continuous clinical oversight, the system analyzes the video feed tracking the patient's hand and toy movements. A fine-tuned Vision-Language Model interprets the recorded video sequences and generates semantic descriptions of the observed activities, enabling high-level verification of the executed actions with respect to the initial textual instructions. A dedicated dataset of 4634 tabletop cognitive remediation video scenarios was collected to evaluate the proposed approach. Experimental results demonstrate that our specialized framework effectively bridges low-level physical telemetry with high-level clinical feedback, presenting a scalable and objective solution for advanced cognitive rehabilitation.