NTSEBENCH: Cognitive Reasoning Benchmark for Vision Language Models

📅 2024-07-15
🏛️ arXiv.org
📈 Citations: 1
Influential: 0
📄 PDF
🤖 AI Summary
Existing benchmarks inadequately evaluate large models’ deep cognitive reasoning—particularly non-commonsense logic, analogy, and sequential reasoning—especially for vision-language models (VLMs). Method: We introduce NTSEBench, the first cognitive reasoning benchmark for VLMs, comprising 2,728 image-text multiple-choice questions derived from India’s National Talent Search Examination (NTSE), covering 26 abstract reasoning categories. We propose four novel modality-cooperative modeling strategies enabling unified evaluation of both open- and closed-source models, alongside cross-modal alignment prompting and a structured annotation framework. Results: Systematic evaluation of state-of-the-art VLMs and LLMs reveals critical bottlenecks in abstract pattern recognition and joint spatial-semantic reasoning. NTSEBench provides a reproducible, high-discrimination, standardized benchmark for advancing cognitive AI research.

Technology Category

Application Category

📝 Abstract
Cognitive textual and visual reasoning tasks, including puzzles, series, and analogies, demand the ability to quickly reason, decipher, and evaluate patterns both textually and spatially. Due to extensive training on vast amounts of human-curated data, LLMs and VLMs excel in common-sense reasoning tasks, however still struggle with more complex reasoning that demands deeper cognitive understanding. We introduce NTSEBench, a new dataset designed to evaluate cognitive multi-modal reasoning and problem-solving skills of large models. The dataset contains 2728 multiple-choice questions, accompanied by a total of 4,642 images, categorized into 26 different types. These questions are drawn from the nationwide NTSE examination in India and feature a mix of visual and textual general aptitude challenges, designed to assess intelligence and critical thinking skills beyond mere rote learning. We establish baselines on the dataset using state-of-the-art LLMs and VLMs. To facilitate a comparison between open source and propriety models, we propose four distinct modeling strategies to handle different modalities -- text and images -- in the dataset instances.
Problem

Research questions and friction points this paper is trying to address.

Large Language Models
Multimodal Models
Complex Reasoning
Innovation

Methods, ideas, or system contributions that make the work stand out.

NTSEBench
Multimodal Understanding
Evaluation Framework