LifePlanner: Evaluating LLM Agents for Geo-spatial Planning with Social Media Data

📅 2026-08-25
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过LifePlanner评估大语言模型在地理空间规划中的表现,该平台结合了社交媒体数据和地图信息,揭示现有模型在复杂任务上的不足。
📝 Abstract
Geo-spatial planning, like trip design, is a realistic testbed for LLM agents because it requires grounded tool use, noisy evidence retrieval, and multi-constraint reasoning. Most benchmarks, however, only provide clean geospatial data and tools, missing the open-ended social signals that people use in daily planning. We introduce LifePlanner, a benchmark that enriches map data with large-scale local social media posts and provides access through an MCP toolset. LifePlanner provides an evaluation suite spanning four task categories and three difficulty levels. Experiments show frontier LLMs perform well on simple retrieval but degrade sharply on complex planning, with the Pass Rate dropping to 40.2%. Results show that failures mainly stem from incomplete evidence acquisition from such a large multimodal database, imprecise tool use, and weak constraint integration rather than model size or reasoning length, suggesting that future progress requires effective grounded planning instead of scaling alone.
Problem

Research questions and friction points this paper is trying to address.

Geo-spatial Planning
Social Media Data
LLM Agents
Complex Planning
Evidence Acquisition
Innovation

Methods, ideas, or system contributions that make the work stand out.

Geo-spatial Planning
Social Media Data
MCP Toolset
Multi-constraint Reasoning
Grounded Planning
🔎 Similar Papers
No similar papers found.