PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the lack of resilience evaluation for self-evolving agents in dynamic environments by proposing a simulator-based code evolution benchmark comprising 144 source-target environment pairs and a diagnostic sandbox feedback mechanism. Evaluating ten methods reveals that the best-performing model solves only 66.7% of static tasks, demonstrating that simulator-driven feedback significantly outperforms unverified self-correction and that parameter inference is not the primary bottleneck. Crucially, the findings identify mechanism redesign as the core obstacle to dynamic adaptation. Consequently, this work provides a critical benchmark and methodological foundation for enhancing agents' capabilities in code iteration and recovery within mutating environments, thereby advancing the assessment of adaptive intelligence under environmental uncertainty.
📝 Abstract
Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchmark of 144 source-to-target adaptation pairs across six physics domains. Each pair links a source environment to a mutated target environment with the same goal and interface. A code-driven design that succeeds in the source fails in the target, where agents must iteratively adapt it into a working target design using diagnostic sandbox feedback within a limited attempt budget. We compare ten self-evolving methods from four paradigms. The benchmark remains far from saturated: Reflexion + Qwen3-14B succeeds on only 35.9\% of full-benchmark pairs, while GPT-5.5 solves 66.7\% of the Statics subset under the full budget. Together, these results show that simulator-grounded reflection is more reliable than unverified self-revision, while memory anchors agents to early designs and broad tree search explores without converging. Even revealing exact physical changes does not raise the performance ceiling, pointing to mechanism redesign rather than parameter inference as the central bottleneck. Data and code are available at https://github.com/thunlp/PACE-Bench.
Problem

Research questions and friction points this paper is trying to address.

Self-evolving agents
Physics adaptation
Dynamic environments
Benchmark
Code evolution
Innovation

Methods, ideas, or system contributions that make the work stand out.

Physics Adaptation
Code Evolution
Self-evolving Agents
Simulator-grounded Reflection
Dynamic Environments
🔎 Similar Papers