CARE: Controlling LLM-Generated Policies through Auditable Review of Evidence in Scientific Experimentation
This work addresses the challenge of safely harnessing large language models (LLMs) in high-throughput experimental optimization, where direct LLM use risks unsafe exploration yet complete exclusion forfeits their optimization potential. To reconcile this trade-off, the authors propose the CARE framework, which employs a non-LLM default optimizer as the primary pathway while leveraging the LLM to generate candidate strategies. Adoption of these candidates is governed by an evidence-based intervention gating mechanism that audits proposals against publicly available evidence, ensuring decisions are auditable, controllable, and traceable. By synergistically integrating LLM-driven creativity with evidence-guided safety constraints, CARE achieves state-of-the-art performance on the Minerva/Olympus and ChemLex benchmarks, improving peak scores from 80.0 to 88.5 and from 83.9 to 92.1, respectively.