GameXpert-Bench: How Far Are Coding Agents from Expert Game Development?

📅 2026-08-22
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
论文提出GameXpert-Bench,通过三个阶段(游戏生成、修复和优化)评估大语言模型作为编码代理在游戏开发中的表现。
📝 Abstract
Recent large language models (LLMs) can operate as coding agents that build complete games from natural language requests. Game development is especially demanding because program logic, visual and audio content, interfaces, interaction and playability must function together in one executable artifact. Measuring this capability therefore requires evaluation of both game product and the development process. Existing benchmarks often assess the game development capabilities of LLMs by evaluating the final artifact or an isolated development stage. Our analysis of complete human-agent development trajectories identifies three stages that together span the lifecycle of game development with a coding agent: initial game generation, bug diagnosis and repair, and optimization over multiple turns. Therefore, we introduce GameXpert-Bench, which operationalizes the three lifecycle stages as three complementary benchmark tracks. GameGen evaluates complete game creation from a single request in an empty workspace. GameFix evaluates diagnosis and repair when defects are reported or left for the agent to discover. GameOpt evaluates cumulative optimization through request chains seeded by real development trajectories between users and agents. We evaluate each track using live game interaction, deterministic behavioral tests, or final product criteria with regression checks. The suite contains 97 generation tasks across 11 genres; 100 repair tasks from 50 game levels verified by humans, each with 19-27 injected bugs; and 17 optimization chains with six turns and 102 requests. Across the three tracks, current agents are more reliable at producing playable foundations and implementing explicit requirements than at discovering defects, verifying runtime behavior, and preserving functionality across changes.
Problem

Research questions and friction points this paper is trying to address.

large language models
game development
coding agents
benchmark
development lifecycle
Innovation

Methods, ideas, or system contributions that make the work stand out.

GameXpert-Bench
coding agents
game development lifecycle
benchmark tracks
evaluation criteria
🔎 Similar Papers
K
Kun Chen
Hunyuan Team, Tencent
H
Haorong Hong
Lightspeed Studios, Tencent
P
Peizhong Gao
Hunyuan Team, Tencent
Jianfeng Lin
Jianfeng Lin
Tsinghua University
Gauge theorylow dimensional topology
T
Tongxu Luo
Hunyuan Team, Tencent
Yuxuan Xie
Yuxuan Xie
Tencent
Reinforcement LearningMulti-Agent Reinforcement Learning
C
Chenxu Liu
Hunyuan Team, Tencent
J
Jieling He
Lightspeed Studios, Tencent
Zhongyuan Liu
Zhongyuan Liu
Tencent
AIGC Games
Z
Zeno Zeng
Hunyuan Team, Tencent