An Interactive Vision Language Platform for Cognitive Remediation in Schizophrenia

πŸ“… 2026-07-21
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
This study addresses the limitations of manual, observer-based assessment of patient actions in cognitive rehabilitation for schizophrenia, which suffers from high subjectivity and poor scalability. To overcome these challenges, this work proposes the first automated evaluation platform leveraging a fine-tuned vision-language model. Operating within a customized tabletop miniature environment, the system integrates video recordings with audio instructions to enable end-to-end mapping from hand–object motion trajectories to high-level clinical semantics. By analyzing video sequences, tracking movements, and generating semantic interpretations, the framework automatically validates the correctness of goal-directed behaviors. Evaluated on a dataset comprising 4,634 videos, the method achieves high-precision semantic-level assessment, substantially enhancing objectivity and scalability, and offers an innovative automated solution for cognitive rehabilitation.
πŸ“ Abstract
Cognitive remediation tasks often require patients to perform structured actions involving object manipulation and sequential reasoning. For patients diagnosed with schizophrenia, these tasks are crucial for addressing severe cognitive deficits. However, evaluating the correctness of these physical actions generally relies on manual observation by clinicians, which introduces subjectivity and limits the scalability of therapeutic interventions. In this paper, we propose an automated framework based on Vision-Language Models for action verification in cognitive remediation tasks tailored for schizophrenia rehabilitation. The proposed system relies on a camera-monitored tabletop environment composed of structured miniature scenes including roads, a roundabout, a park, and toy vehicles. Patients receive audio instructions describing goal-oriented spatial actions to perform by manipulating a toy vehicle. These interactive physical activities are specifically designed to stimulate targeted cognitive functions, such as sustained attention, motor coordination, spatial navigation, and cognitive flexibility. To verify the correctness of the performed actions without requiring continuous clinical oversight, the system analyzes the video feed tracking the patient's hand and toy movements. A fine-tuned Vision-Language Model interprets the recorded video sequences and generates semantic descriptions of the observed activities, enabling high-level verification of the executed actions with respect to the initial textual instructions. A dedicated dataset of 4634 tabletop cognitive remediation video scenarios was collected to evaluate the proposed approach. Experimental results demonstrate that our specialized framework effectively bridges low-level physical telemetry with high-level clinical feedback, presenting a scalable and objective solution for advanced cognitive rehabilitation.
Problem

Research questions and friction points this paper is trying to address.

cognitive remediation
schizophrenia
action verification
vision-language models
automated assessment
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Cognitive Remediation
Action Verification
Automated Assessment
Schizophrenia Rehabilitation
N
Nassira Ait Mehdi
RIIMA Laboratory, Computer Science Faculty, USTHB University, 16111, Algeria
M
Milissa Temmam
RIIMA Laboratory, Computer Science Faculty, USTHB University, 16111, Algeria
Slimane Larabi
Slimane Larabi
Faculty of Computer Science, University of Science and Technology Houari Boumediene
Artificial IntelligenceComputer visionAccessibilityBrain Computer Interaction