Anchored Scenario Coverage for Failure-Aware First-Hit Batch Inverse Design
研究提出ARC-SC方法,通过保留强边际候选作为锚点并最大化预测目标场景的互补覆盖率,以提高在实验失败情况下的早期有效设计发现效率。
研究提出ARC-SC方法,通过保留强边际候选作为锚点并最大化预测目标场景的互补覆盖率,以提高在实验失败情况下的早期有效设计发现效率。
This work addresses the lack of systematic evaluation frameworks for inverse design algorithms in materials science, as existing machine learning benchmarks are largely confined to forward property prediction. To bridge this gap, we introduce MatFormBench—the first unified benchmark for goal-driven materials formulation. Built upon a physics-informed synthetic data generation pipeline, MatFormBench features five tiers of task difficulty and a multidimensional scoring metric, MatFormScore, which evaluates performance across target achievement, search efficiency, exploration capability, robustness, and stability. Through standardized evaluations of 39 algorithms—including diffusion models, variational autoencoders (VAEs), genetic algorithms, and large language models—across 1,170 trials, we demonstrate MatFormBench’s effectiveness: diffusion models emerge as overall top performers, while VAEs and genetic algorithms excel in specific scenarios, underscoring the benchmark’s value in algorithm assessment, diagnostic analysis, and reproducibility.
研究提出ARC-SC方法,通过保留强边际候选作为锚点并最大化预测目标场景的互补覆盖率,以提高在实验失败情况下的早期有效设计发现效率。
This work addresses the lack of systematic evaluation frameworks for inverse design algorithms in materials science, as existing machine learning benchmarks are largely confined to forward property prediction. To bridge this gap, we introduce MatFormBench—the first unified benchmark for goal-driven materials formulation. Built upon a physics-informed synthetic data generation pipeline, MatFormBench features five tiers of task difficulty and a multidimensional scoring metric, MatFormScore, which evaluates performance across target achievement, search efficiency, exploration capability, robustness, and stability. Through standardized evaluations of 39 algorithms—including diffusion models, variational autoencoders (VAEs), genetic algorithms, and large language models—across 1,170 trials, we demonstrate MatFormBench’s effectiveness: diffusion models emerge as overall top performers, while VAEs and genetic algorithms excel in specific scenarios, underscoring the benchmark’s value in algorithm assessment, diagnostic analysis, and reproducibility.