CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过创建CulturalMenuBench基准来探讨多模态语言模型在食物识别与文化理解间的知识应用差距,揭示了模型虽能高精度识别图像但缺乏将视觉信息与文化背景结合的能力。
📝 Abstract
Multimodal language models achieve near-ceiling scores on food recognition benchmarks, yet it remains unclear whether this success reflects genuine cultural understanding or mere visual matching. To probe this distinction, we introduce CulturalMenuBench, a benchmark of 4,870 items in 10 languages across 18 regions; its 10 tasks pair final-dish and step-by-step cooking images with ingredients, procedural text, and regional labels, spanning basic recognition to process-grounded cultural attribution. Evaluating 12 models exposes a substantial knowledge-application gap: models exceeding 94% on standard multiple-choice tasks drop to at most 56% when attributing dishes to Chinese regional cuisines, despite an identical four-way format. Diagnostic analyses explain why: error patterns are consistent with random guessing, accuracy tracks visual distinctiveness rather than cultural structure, and models classify cuisines more accurately from dish names alone than from images (+7-18 points). The knowledge is thus present but cannot be activated through visual input. An ablation confirms these tasks genuinely require procedural evidence: removing sequential cooking images selectively degrades process-grounded tasks while others remain stable. Overall, CulturalMenuBench shows that near-perfect recognition can conceal an inability to apply cultural knowledge, motivating training that explicitly connects perception, procedure, and cultural context. Code and data are publicly available.
Problem

Research questions and friction points this paper is trying to address.

multimodal language models
cultural understanding
visual matching
knowledge-application gap
cultural context
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cultural Understanding
Multimodal Reasoning
Knowledge-Application Gap
Culinary Benchmark