A Judge Agent Closes the Reliability Gap in AI-Generated Scientific Simulation

📅 2026-03-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the high prevalence of silent failures in AI-generated scientific simulation code when applied to non-textbook problems, which undermines its reliability. To bridge this trust gap, we propose the Judge Agent framework, which systematically validates well-posedness, convergence, and error bounds through automated verification. We introduce a simulatability class \( \mathcal{S} \) and a structured, solver-agnostic specification format (spec.md) that enables machine-readable problem descriptions. Evaluated across 134 cases spanning 12 scientific domains, our approach reduces the silent failure rate from 42% to 1.5%. In 72 blind test tasks, it achieves an 89% success rate, and in clinical CT reconstruction experiments, it reaches 99% of expert-level performance, substantially enhancing the credibility of AI-generated scientific code.

Technology Category

Application Category

📝 Abstract
Large language models can generate scientific simulation code, but the generated code silently fails on most non-textbook problems. We show that classical mathematical validation -- well-posedness, convergence, and error certification -- can be fully automated by a Judge Agent, reducing the silent-failure rate from 42% to 1.5% across 134 test cases spanning 12 scientific domains. The headline result comes from a prospective benchmark: 72 blinded tasks submitted by 12 independent scientists yield an 89% success rate (95% CI: [80%, 95%]) with automated error bounds, versus 53% without the Judge. On clinical CT (the only powered experiment, n = 200), the pipeline reaches 99% of expert quality. The residual 1.5% concentrates at bifurcation points where certifiability breaks down. We formalize this boundary through the simulability class S and introduce spec.md, a structured specification format that makes any scientific computation problem machine-readable and solver-independent. Code, data, and all 72 benchmark tasks are publicly archived.
Problem

Research questions and friction points this paper is trying to address.

silent failure
scientific simulation
reliability gap
code generation
validation
Innovation

Methods, ideas, or system contributions that make the work stand out.

Judge Agent
automated validation
simulability class
spec.md
scientific simulation
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
C
Chengshuai Yang
NextGen PlatformAI C Corp, USA