🤖 AI Summary
This study addresses the challenge of balancing factual consistency with meaningful progression in LLM-generated narratives within open-world simulations. We introduce WSE-bench and a novel process-oriented evaluation framework that quantifies dynamic narratives across generation coverage, consistency, and richness dimensions. Our analysis reveals distinct impacts of model scale and architecture, demonstrating that sustained generation, normative consistency, and meaningful development constitute independent yet competing capability dimensions. Crucially, we identify a non-concave Pareto frontier between consistency and richness. These findings provide essential theoretical grounding and standardized benchmarks for understanding and optimizing long-horizon narrative capabilities in large language models, establishing that improving one dimension often entails trade-offs in others within open-ended simulation environments.
📝 Abstract
Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.