AI Research Preference Models

📅 2026-08-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the challenge of high evaluation costs hindering experimental progress in AI research agents by proposing a preference prediction method based on frozen pretrained language models. Leveraging dual-mode reasoning and proxy modeling integrated with the AIRA-dojo framework for optimized budget allocation, this approach enables training-free value assessment without task-specific fine-tuning. Experimental results demonstrate that the proposed method achieves a normalized score of 0.729 while reducing evaluation budgets by over one-third. Furthermore, it attains state-of-the-art performance across two distinct tasks, significantly enhancing both research efficiency and resource utilization in AI scientific discovery.
📝 Abstract
AI research agents (AIRA) can now propose, implement, and evaluate their own machine learning experiments, but progress on frontier tasks is throttled by cost: a candidate solution can be written in minutes, whereas evaluating it can take hours to days of GPU time. An agent can therefore propose far more candidates than it can afford to run, and its progress depends on its research preference: how it allocates a fixed execution budget across many candidates. We introduce AI Research Preference Models (RPMs) that predict which of multiple candidate solutions are most worth executing, without paying the cost of executing them all. We build RPMs from frozen pretrained language models (with no task-specific training), in two forms: an inference-only model that reasons over candidate plans, code, and prior executed solutions, and an agentic model that additionally runs small-scale pilot experiments before deciding. We integrate both into the AIRA-dojo search agent and evaluate on AIRS-Bench, a recent benchmark of machine learning research tasks for AI research agents. The two variants raise the average normalized score from 0.684 to 0.711 and 0.729 respectively, and reach the unguided agent's 24-hour performance in roughly 15 hours, using less than two-thirds of its execution budget. Our best RPMs also yield new state-of-the-art results on two AIRS-Bench tasks.
Problem

Research questions and friction points this paper is trying to address.

AI Research Agents
Research Preference Models
Execution Budget Allocation
Candidate Selection
Evaluation Cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

AI Research Preference Models
Frozen Pretrained Language Models
Budget-aware Search
Pilot Experiments
AIRS-Bench
T
Thomas Simon Foster
FAIR at Meta, University of Oxford
Bassel Al Omari
Bassel Al Omari
University of Waterloo
Tingchen Fu
Tingchen Fu
Renmin University of China
natural language processing
T
Thomas Mann
FAIR at Meta
C
Carl Domond
FAIR at Meta
L
Lucia Cipolina-Kun
FAIR at Meta
B
Bhavul Gauri
FAIR at Meta
M
Muna Aghamelu
FAIR at Meta
Alexander D. Goldie
Alexander D. Goldie
Oxford University
reinforcement learningmeta learningmachine learning
Eryk Helenowski
Eryk Helenowski
Machine Learning Engineer, Meta
LLMAI
J
Jean-Christophe Gagnon-Audet
FAIR at Meta
Alberto Pepe
Alberto Pepe
Sage Bionetworks
Computational AstrophysicsInformation ScienceScholarly Communication
S
Saba Nazir
FAIR at Meta
D
Daniel Izcovich
FAIR at Meta
Noam Levi
Noam Levi
Postdoctoral Fellow, AI4Science/AI Center, EPFL
RMTStatistical Learning TheoryField TheoryParticle Physics
R
Rishi Hazra
FAIR at Meta
Karen Hambardzumyan
Karen Hambardzumyan
FAIR, Meta + University College London
InterpretabilityNatural Language ProcessingFew-Shot Learning
N
Nicolas Baldwin
FAIR at Meta
Xian Li
Xian Li
FAIR, Meta
machine learningnatural language processingmachine translationdeep learning
Martin Josifoski
Martin Josifoski
Meta
Paris Giampouras
Paris Giampouras
Assistant Professor @ University of Warwick
generative modelsrepresentation learningadversarial robustnesscontinual learning
Masoud Jalili Sabet
Masoud Jalili Sabet
FAIR at Meta
Anya Sims
Anya Sims
University of Oxford
Reinforcement LearningDeep Learning
H
Hela Momand
FAIR at Meta
Tatiana Shavrina
Tatiana Shavrina
Meta
Natural language processingcomputational linguisticsbenchmarkingmultilinguality