Autonomous Microscopy Experiments through Large Language Model Agents
Existing self-driving laboratories (SDLs) rely on static experimental protocols, limiting their ability to emulate scientists’ adaptive reasoning and intuition in dynamic environments. Method: We propose AILA, the first large language model (LLM)-based autonomous agent system for end-to-end atomic force microscopy (AFM) experimentation—encompassing experimental design, execution, analysis, and closed-loop decision-making. Contribution/Results: We introduce AFMBench, the first benchmark for evaluating LLMs in AFM-driven scientific discovery, uncovering critical deficiencies in multi-agent coordination (73% failure rate), instruction following, and safety alignment, while empirically delineating LLMs’ scientific reasoning boundaries. Leveraging task-decomposition prompting, hardware interface integration, and a multi-agent architecture, AILA achieves autonomous AFM calibration, high-resolution feature identification, and nanomechanical property quantification. Results further reveal substantial accuracy degradation in foundational tasks (e.g., document retrieval), underscoring robustness and trustworthiness as central challenges in AI for Science.