Efficient Test-Time Adaptation through Human-AI Interaction

📅 2026-09-03
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出通过人机交互进行测试时适应(TAHI),以解决AI生成内容不符合个人专业标准的问题,提升个体任务成功率4.5-20.9%。
📝 Abstract
AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.
Problem

Research questions and friction points this paper is trying to address.

Human-AI Interaction
Test-Time Adaptation
Individual Expertise
Innovation

Methods, ideas, or system contributions that make the work stand out.

Test-Time Adaptation
Human-Agent Interaction
Evolving Rubric Module
Zora Zhiruo Wang
Zora Zhiruo Wang
Language Technologies Institute, Carnegie Mellon University
natural language processing
Apurva Gandhi
Apurva Gandhi
Carnegie Mellon University
Machine LearningArtificial Intelligence
Rulin Shao
Rulin Shao
University of Washington
machine learning
A
Aspen Chen
Stanford University
Jonas Mueller
Jonas Mueller
Cleanlab
Trustworthy AIMachine LearningStatisticsComputational Biology
Z
Zhiqi Liang
University of California San Diego
J
Jett Chen
Carnegie Mellon University
M
Michael Ryan
University of Washington
Q
Qianou Ma
Handshake AI
Luxi He
Luxi He
Department of Computer Science, Princeton University
Zhoujun Cheng
Zhoujun Cheng
UC San Diego
Natural Language ProcessingArtificial Intelligence
Andre He
Andre He
Undergraduate Student, UC Berkeley
Computer Science
Seungone Kim
Seungone Kim
Carnegie Mellon University
Large Language ModelsNatural Language Processing
Jiayi Geng
Jiayi Geng
Carnegie Mellon University
Natural Language ProcessingMachine LearningCognitive Science
Mingqian Zheng
Mingqian Zheng
Language Technologies Institute, Carnegie Mellon University
Natural Language Processing
W
Weiwei Sun
Stanford University
Z
Zheyuan Zhang
Princeton University
X
Xinran Zhao
University of California San Diego
Yike Wang
Yike Wang
University of Washington
Natural Language Processing
A
Abe Hou
University of Washington
L
Liwei Jiang
Handshake AI
Pang Wei Koh
Pang Wei Koh
University of Washington; Allen Institute for AI
Machine learningNatural language processingComputational biology
Diyi Yang
Diyi Yang
Stanford University
Computational Social ScienceNatural Language ProcessingMachine Learning
Graham Neubig
Graham Neubig
Carnegie Mellon University, All Hands AI
Natural Language ProcessingMachine LearningArtificial Intelligence
Daniel Fried
Daniel Fried
Carnegie Mellon University
Natural Language ProcessingMachine Learning