Boosting Knowledge-based Visual Question Answering with Structured Context Reasoning

📅 2026-08-17
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出SCoRe框架,通过结构化上下文推理解决知识型视觉问答中因直接拼接异构信息导致的性能下降问题。
📝 Abstract
Knowledge-based Visual Question Answering aims to answer questions about an image by integrating external knowledge with visual and textual information. Recent approaches often rely on in-context learning to prompt Large Language Models (LLMs) with multimodal context in a zero-shot or few-shot manner. However, we observe that directly concatenating heterogeneous visual descriptions and retrieved knowledge into long, unstructured prompts often degrades reasoning performance, due to both excessive irrelevant context and the lack of explicit relational structure. In this paper, we propose an LLM-based Structured Context Reasoning (SCoRe) framework that infers both explicit and implicit relationships for prediction. SCoRe consists of three stages: Context Acquisition, which generates diverse visual notes and retrieves explicit knowledge via an efficient two-stage multimodal retrieval strategy; Context Selection, which filters relevant visual, explicit, and implicit knowledge using LLM-guided selection; and Context Compression, which performs Relational Logic Distillation (RLD) to transform raw text into explicit entity-relation triplets. These relational triplets serve as a concise and structured prompt for final answer prediction. Extensive experiments on the OK-VQA and A-OKVQA benchmarks demonstrate that SCoRe consistently outperforms state-of-the-art methods.
Problem

Research questions and friction points this paper is trying to address.

Knowledge-based Visual Question Answering
unstructured prompts
reasoning performance
Innovation

Methods, ideas, or system contributions that make the work stand out.

Structured Context Reasoning
Relational Logic Distillation
Large Language Models
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Q
Qiyou Liu
School of Computer Science, Guangdong University of Technology, Guangzhou, China
Yong Zhang
Yong Zhang
The Chinese University of Hong Kong, Shenzhen
Vision-language multimodal learningAI
J
Jianjie Luo
School of Computer Science, Guangdong University of Technology, Guangzhou, China
Z
Zhenguo Yang
School of Computer Science, Guangdong University of Technology, Guangzhou, China
Yi Yu
Yi Yu
Graduate School of Advanced Science and Engineering at Hiroshima University
Multimodal learningGenerative modelingMultimediaAI Music