Assisted Spatial Cognition Through Vision-Language Models

📅 2026-09-11
📈 Citations: 0
✹ Influential: 0
📄 PDF
🀖 AI Summary
本文提出䞀种结合倧型语蚀暡型、视觉-语蚀暡型和数字孪生技术的新框架以增区视障及神经倚样性人士的空闎讀知富航胜力通过手机摄像倎捕捉视频并生成3D场景理解。
📝 Abstract
Multimodal AI, powered by Large Language Models (LLMs) and Vision-Language Models (VLMs), is transforming assistive technologies by enabling simultaneous processing of visual and textual data. This advancement holds significant promise for over 43 million visually impaired and neuro-divergent individuals worldwide who face persistent challenges in navigating indoor and outdoor environments due to limited spatial awareness and insufficient environmental cues. Existing navigation aids often lack comprehensive 3D scene understanding, relying on constrained route-based strategies that hinder user autonomy. In this paper, we introduce a novel end-to-end framework that integrates LLMs, VLMs and digital twin technologies to deliver a spatially cognitive navigation support for visually impaired and neuro-divergent users. Our system captures video input via standard mobile phone cameras, and employs SLAM3R to generate dense 3D point clouds from monocular RGB sequences in real-time. Our custom post-processing algorithm ensures accurate point cloud alignment across multiple viewpoints without requiring predefined reference points. This enhances the capabilities of SpatialLM to produce structured 3D representations, including architectural elements and oriented object bounding boxes. The enriched spatial data is then processed by a locally deployed LLM, which interprets 3D contexts to generate detailed scene descriptions and precise distance measurements between users and surrounding objects. We evaluated our approach across diverse video scenarios featuring various perspectives, looped walking views and captured in multiple environments. The evaluation results demonstrate consistent accuracy in 3D scene interpretation and object localisation, underscoring the potential of our system as a transformative assistive navigation solution that combines advanced visual perception with spatial reasoning
Problem

Research questions and friction points this paper is trying to address.

Visually Impaired
Neuro-divergent
Navigation Aids
Spatial Awareness
3D Scene Understanding
Innovation

Methods, ideas, or system contributions that make the work stand out.

Vision-Language Models
Spatial Cognition
3D Point Clouds
SLAM3R
Digital Twin Technologies
🔎 Similar Papers
No similar papers found.
H
H. Riaz
School of Electronic Engineering, Dublin City University, Dublin, Ireland; Insight Research Ireland Centre for Data Analytics, Dublin, Ireland
J
J. B. Fernandez
School of Electronic Engineering, Dublin City University, Dublin, Ireland; Insight Research Ireland Centre for Data Analytics, Dublin, Ireland
I
I. Mills
South East Technological University, Waterford, Ireland
D
D. Hickey
South East Technological University, Waterford, Ireland
F
F. Cleary
South East Technological University, Waterford, Ireland
M
M. I. Ali
School of Electronic Engineering, Dublin City University, Dublin, Ireland; Insight Research Ireland Centre for Data Analytics, Dublin, Ireland