Praxist: From Experimental Artifacts to Solution Lineages

📅 2026-08-26
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究解决了自主研发代理在实验中难以追踪改进原因和成本过高的问题,通过引入Praxist系统,将可重复的实验结果转化为证据图谱,以更低的成本获得更优的结果。
📝 Abstract
Autonomous R\&D agents now write, run, and improve executable artifacts under automated evaluation---but largely as laboratory instruments: shown on curated benchmarks, with gains that are hard to trace to a cause and costs well above what sustained engineering practice absorbs. The limitation is structural. Most systems treat each attempt as nearly self-contained, so logs, memories, and search trees record what happened without establishing which design element produced an improvement, whether its evidence survived validation, or how it recombines with others. Long campaigns therefore keep re-learning the same lessons. We introduce Praxist, a lineage-centered generational system that converts reproducible artifacts and evaluator outcomes into a typed evidence graph of findings, lane-structured frontiers, and agendas. Separating local artifact construction from cohort-level evidence synthesis lets later attempts inherit validated mechanisms, unresolved claims, and useful constraints, and leaves results attached to an inspectable lineage. On the standardized 75-task MLE-bench suite, the finalized official-grader results give Praxist 60 medals (80.0\%), 49 of them gold, against 55 medals (73.3\%) and 34 gold for a Claude Code baseline on Claude Opus 4.8---at a recorded model spend of US\$3,054 versus US\$38,370, roughly a twelfth of the cost. Four case studies---quantitative trading, LiDAR-inertial-visual SLAM, tokamak magnetic control, and rocket landing---carry the same process into open-ended engineering problems, improving on each task-native baseline in headline accuracy, survival, or resource cost, with the discovery path on record. Stronger artifacts at an order of magnitude less spend, each backed by an auditable lineage, are, to our knowledge, first brought together here: the operating profile production research requires, not the one a benchmark demonstration establishes.
Problem

Research questions and friction points this paper is trying to address.

autonomous R&D agents
executable artifacts
automated evaluation
evidence tracking
cost efficiency
Innovation

Methods, ideas, or system contributions that make the work stand out.

lineage-centered
evidence graph
cost-effective
artifact construction
cohort-level evidence synthesis
🔎 Similar Papers
💼 Related Jobs
No related jobs found.
J
Jin Li
Sapient Intelligence
A
Ahmed Murtadha
Sapient Intelligence
Z
Zhiyu Wang
Sapient Intelligence
Q
Qiwen Chen
Sapient Intelligence
William Chen
William Chen
Carnegie Mellon University
Spoken Language ProcessingSpeech RecognitionSpeech TranslationMachine Translation
Y
Yifei Wu
Sapient Intelligence
Guan Wang
Guan Wang
CEO, Sapient Intelligence
Artificial General IntelligenceReinforcement LearningLarge Language Models
A
Andy L. Siy
Sapient Intelligence
J
Jiayi Yang
Sapient Intelligence
M
Mengsha Huang
Tsinghua University
W
Wenhao Li
Sapient Intelligence
Yixuan Liu
Yixuan Liu
AMD, Tsinghua University
Generative AI
S
Shuailin Pan
Sapient Intelligence
M
Mingli Yuan
Sapient Intelligence
Sen Song
Sen Song
Laboratory of Brain and Intelligence, Dept of Biomedical Engineering, Tsinghua University
Brain-inspired ComputationComputational NeurocienceArtificial General IntelligenceScience of HappinessNeural Circuits
Y
Yuhao Sun
Sapient Intelligence