Vision And Text Transformer For Predicting Answerability On Visual Question Answering

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文针对视觉问答的可回答性预测问题,提出了一种基于Transformer架构的方法VT-Transformer来处理图像和文本特征,实验表明其有效性。
📝 Abstract
Answerability on Visual Question Answering is a novel and attractive task to predict answerable scores between images and questions in multi-modal data. Existing works often utilize a binary mapping from visual question answering systems into Answerability. It does not reflect the essence of this problem. Together with our consideration of Answerability in a regression task, we propose VT-Transformer, which exploits visual and textual features through Transformer architecture. Experimental results on VizWiz 2020 dataset show the effectiveness and robustness of VT-Transformer for Answerability on Visual Question Answering when comparing with competitive baselines.
Problem

Research questions and friction points this paper is trying to address.

Answerability
Visual Question Answering
regression task
Innovation

Methods, ideas, or system contributions that make the work stand out.

VT-Transformer
Visual and Textual Features
Regression Task
Answerability
🔎 Similar Papers
No similar papers found.