MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft

📅 2026-08-28
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过引入MineCEraft基准,评估语言模型在Minecraft中执行建设任务的能力,使用723条专家级自然语言指令进行系统性测试。
📝 Abstract
We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, spanning 17 distinct task categories, providing a safe and controllable experimental environment for assessing LLMs' ability to perform realistic construction engineering tasks. With this benchmark, we conduct an in-depth evaluation of state-of-the-art LLMs and perform a detailed error analysis, revealing key failure modes and practical challenges in applying LLMs to construction engineering tasks.
Problem

Research questions and friction points this paper is trying to address.

Language Models
Construction Engineering
Minecraft
Benchmark Evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark
Construction Engineering
Natural-Language Instructions
Evaluation
Error Analysis
🔎 Similar Papers
No similar papers found.
S
Sewoong Lee
Siebel School of Computing and Data Science, The Grainger College of Engineering, University of Illinois Urbana-Champaign
R
Risham Sidhu
Siebel School of Computing and Data Science, The Grainger College of Engineering, University of Illinois Urbana-Champaign
Julia Hockenmaier
Julia Hockenmaier
Professor, University of Illinois at Urbana-Champaign (UIUC)
Natural Language ProcessingComputational LinguisticsArtificial Intelligence
Y
Yoonhwa Jung
Department of Civil and Coastal Engineering, University of Florida