Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic Discovery
This work addresses a critical limitation in existing benchmarks for scientific equation discovery, which often fail to distinguish whether models genuinely infer underlying laws or merely reproduce known formulas. To this end, the authors propose the LSR-Synth evaluation framework, which introduces novel synthetic terms into established mechanisms and incorporates strategies such as semantic blinding, operator library weakening, and exclusion of matching operator families to construct a semantics-free baseline. This setup rigorously isolates and quantifies the marginal contribution of language model priors. Experiments reveal that, under current task sets and search budgets, fixed operator libraries already cover most problems; only when lexical coverage is selectively disrupted do language model–generated candidate solutions substantially increase the number of solvable instances, thereby providing the first quantitative evidence of their practical value in out-of-distribution symbolic regression.