Measuring Decision-Scale Use in Tool-Augmented LLMs: A Contrastive Urban Benchmark

📅 2026-09-15
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入URBANCONTRASTIVEQA基准,评估工具增强的语言模型在对比城市活动异常程度时的表现,提出需要展示本地基线而不仅仅是活动量。
📝 Abstract
Urban decision-support often asks whether activity is unusually high or low for a specific place, not which place has the larger raw count. Twenty pickups in a quiet neighborhood can be more abnormal than 180 at an airport. We introduce URBANCONTRASTIVEQA, a benchmark that asks whether tool-augmented language models can make this baseline-relative comparison. Each item pairs two urban situations from public mobility data in NYC, Chicago, and Seattle, labeled by how far current activity deviates from that place's historical baseline. We evaluate six instruction-tuned models under five tool-output formats. With only raw counts, models often pick the larger number even when it is less abnormal for its zone. Server-computed baseline scores and ordinal labels raise accuracy, but gains vary by model. For heterogeneous urban feeds, tool interfaces need to expose local baselines, not just activity volumes. We release the pair bank, labels, scoring scripts, and data card.
Problem

Research questions and friction points this paper is trying to address.

Urban Decision-Support
Baseline-Relative Comparison
Tool-Augmented Language Models
Innovation

Methods, ideas, or system contributions that make the work stand out.

URBANCONTRASTIVEQA
baseline-relative comparison
tool-augmented language models
local baselines
R
Ray Chen
Department of Computer & Information Science Engineering, University of Florida
V
Vivian Wong
College of Design, Construction and Planning, University of Florida
Christan Grant
Christan Grant
Associate Professor, University of Florida
Interactive Machine LearningNatural Language ProcessingVisualizationData MiningPrivacy