Overconfident Oracles: Limitations of In Silico Sequence Design Benchmarking
This paper identifies a critical unreliability in machine learning–based oracle benchmarks for biological sequence design: over 70% of methods exhibit rank reversal across different oracle models—due to architectural or training stochasticity—revealing a fundamental deficiency in out-of-distribution generalization. Method: The authors systematically characterize the detrimental impact of oracle inconsistency on benchmark validity and propose a hybrid evaluation framework integrating multiple biophysically grounded metrics—including stability, foldability, and solubility—while constraining the oracle’s scoring domain to enhance design robustness. Contribution/Results: Through large-scale reproduction of 12 state-of-the-art design methods, cross-oracle consistency analysis, and out-of-distribution generalization diagnostics, the framework significantly improves wet-lab validation rates. It establishes a new paradigm for constructing trustworthy, AI-driven benchmarks for protein and nucleic acid design—grounded in empirical feasibility and physical realism.