When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations

📅 2026-08-16
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of balancing factual consistency with meaningful progression in LLM-generated narratives within open-world simulations. We introduce WSE-bench and a novel process-oriented evaluation framework that quantifies dynamic narratives across generation coverage, consistency, and richness dimensions. Our analysis reveals distinct impacts of model scale and architecture, demonstrating that sustained generation, normative consistency, and meaningful development constitute independent yet competing capability dimensions. Crucially, we identify a non-concave Pareto frontier between consistency and richness. These findings provide essential theoretical grounding and standardized benchmarks for understanding and optimizing long-horizon narrative capabilities in large language models, establishing that improving one dimension often entails trade-offs in others within open-ended simulation environments.
📝 Abstract
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.
Problem

Research questions and friction points this paper is trying to address.

Open-ended storytelling
World simulations
Narrative consistency
LLM evaluation
Dynamic storytelling
Innovation

Methods, ideas, or system contributions that make the work stand out.

WSE-bench
Process Benchmark
Open-Ended Storytelling
Pareto Frontier
Narrative Consistency
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
Yuqi Chen
Yuqi Chen
Fudan University, Xiaohongshu Inc.
Maching LearningTime SeriesRecommendation System
S
Sixuan Li
Peking University
Yunfeng Cai
Yunfeng Cai
Beijing Institute of Mathematical Sciences and Applications (BIMSA)
AI
X
Xueai Li
The University of Hong Kong
K
Ka Man Yan
The University of Hong Kong
Y
Ying Li
Tsinghua University