DiG-bench: Discovery in Games

πŸ“… 2026-08-12
πŸ“ˆ Citations: 0
✨ Influential: 0
πŸ“„ PDF
πŸ€– AI Summary
Existing AI benchmarks struggle to evaluate agents’ capacity to discover novel knowledge through experimentation in environments with unknown objectives. To address this gap, this work proposes a benchmark platform comprising 70 interactive games, each implicitly encoding a unique transformation rule via a short string and featuring an unknown win condition. The platform operationalizes the scientific discovery skill of formulating new generalizations into measurable tasks for the first time, incorporating a seven-level difficulty gradient that remains solvable by humans yet challenging for AI systems. All games have been solved by at least one human on their first attempt; while lower-difficulty levels are routinely mastered by various models, higher-difficulty levels continue to pose significant challenges even for state-of-the-art agents. Twenty-one of these games have been publicly released.
πŸ“ Abstract
Discovery---formulating novel generalizations---is a central part of the scientific process. Despite its importance, there is a gap in the current AI benchmark landscape, with few benchmarks directly probing the capacity for discovering new knowledge with experimentation in controlled environments where the objective is unknown. To address this gap, we release a new benchmark: DiG-bench (Discovery in Games). DiG-bench consists of a set of 70 independent games. Each game is encoded as a short string and has unique transformation rules that must be discovered through interaction and experimentation. The levels of the game present a series of challenges to test whether the rules have been discovered, where the win conditions for each level are also unknown. We provide games at seven tiers of difficulty for AI agents. The lowest tier is routinely solvable by multiple models, while the highest tier challenges the best models in agentic harnesses. All 70 games were solved by at least one human on first attempt. A subset of 21 games is released publicly, and the remainder is held private for secure evaluation.
Problem

Research questions and friction points this paper is trying to address.

discovery
AI benchmark
unknown objective
experimentation
generalization
Innovation

Methods, ideas, or system contributions that make the work stand out.

discovery
benchmark
game-based evaluation
unknown objectives
AI experimentation
πŸ”Ž Similar Papers
No similar papers found.
R
Ruairidh M. Battleday
Thinking About Thinking
K
Kai Sandbrink
Independent
J
Jimi Cullen-Drohan
Independent
Zihan Yan
Zihan Yan
University of Illinois Urbana-Champaign
PsychometricsNLP/MLSensing
T
Timothy Muller
University of Oxford
C
Clare Maguire
Thinking About Thinking
A
Ales Kubicek
Independent
F
Fraser Greenlee-Scott
Independent
S
Sukrit Sumant
Independent
Tri Dao
Tri Dao
Princeton University, Together AI
Machine learningSystems
J
JΓΌrgen Schmidhuber
King Abdullah University of Science and Technology, Swiss AI Lab
Michal Valko
Michal Valko
Chief Models Officer @ Stealth Startup, Inria & MVA - Ex: Llama at Meta; Gemini and BYOL @ Deepmind
large language modelsreasoningfine-tuningtest-time computationRLHF
J
Joshua Tenenbaum
Massachusetts Institute of Technology
Thomas L. Griffiths
Thomas L. Griffiths
Professor of Psychology and Computer Science, Princeton University
Computational Models of CognitionCognitive ScienceMachine LearningCognitive PsychologyBayesian Statistics
Z
Zeb Kurth-Nelson
Independent
J
James C. R. Whittington
Thinking About Thinking, University of Oxford