Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

📅 2026-09-04
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文探讨了将常识知识融入计算机视觉以提高场景理解和物体间关系推理的问题,通过知识图谱、场景图等方法,并指出了当前局限性和未来研究方向。
📝 Abstract
Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning models but also enhances their ability to interact meaningfully with humans and the environment. Unlike CNN-based conventional vision models, which are designed to identify objects within a specific image, incorporating commonsense knowledge enables models to interpret scenes in a more holistic manner, thereby improving their spatial ability to reason about relationships among objects and actions. This integration not only enhances object recognition but also facilitates a deeper understanding of the contextual factors, ultimately leading to more precise predictions and interactions in real-world applications. This paper presents a comprehensive survey of recent developments that integrate commonsense knowledge into computer vision tasks. We systematically review approaches based on knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. We also outline current limitations related to dataset bias, knowledge incompleteness, and integration challenges. Finally, we highlight prospective research trajectories in cross-modal reasoning, scalable commonsense knowledge injection, and neuro-symbolic hybrid architectures to develop truly intelligent visual systems.
Problem

Research questions and friction points this paper is trying to address.

Commonsense Reasoning
Computer Vision
Contextual Knowledge
Scene Understanding
Object Recognition
Innovation

Methods, ideas, or system contributions that make the work stand out.

knowledge graphs
scene graphs
neuro-symbolic models
commonsense-augmented transformers
💼 Related Jobs
No related jobs found.
B
Bahar Uddin Mahmud
Lander University, USA
S
Sumit Barua
Western Michigan University, USA
G
Guan Yue Hong
Western Michigan University, USA
Ajay Gupta
Ajay Gupta
Western Michigan University, USA
H
Hexu Liu
Western Michigan University, USA