Are We There Yet? Assessing Computer-Use Agents for Blind Users' Accessible Interaction with Desktop Applications

📅 2026-08-31
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过三周日记研究评估了计算机使用代理对盲人用户访问桌面应用程序的有效性,使用OLLA原型收集数据,并比较了GPT-5等五种模型的表现。
📝 Abstract
Computer-use agents are emerging as a paradigm for agentic human-AI interaction, combining language reasoning with multi-modal interface grounding to operate GUIs. Yet their effectiveness for blind screen-reader users in real-world desktop workflows remains unclear. We present a three-week diary study with 8 blind users using OLLA, a screen-reader-accessible CUA prototype, collecting 1,258 commands across 12 applications with screenshots, UI trees, model responses, and action traces. We evaluate GPT-5 during deployment and re-execute the same commands with four additional models. GPT-5 achieved the highest success rate at 52.5%. Trace analysis reveals grounding, planning, constraint-tracking, and termination failures, while interviews reveal beyond-automation needs.
Problem

Research questions and friction points this paper is trying to address.

Computer-use agents
Blind users
Accessible interaction
Desktop applications
Screen-reader
Innovation

Methods, ideas, or system contributions that make the work stand out.

Computer-use agents
Screen-reader users
GUI operation
GPT-5
Trace analysis
🔎 Similar Papers
2024-03-22International Conference on Human Factors in Computing SystemsCitations: 16
💼 Related Jobs
No related jobs found.
S
Satwik Ram Kodandaram
Stony Brook University
M
Monalika Padma Reddy
Stony Brook University
Xiaojun Bi
Xiaojun Bi
Department of Computer Science, Stony Brook University
Human Computer InteractionMobile User InterfacesText InputHuman Performance Models
Jiawei Zhou
Jiawei Zhou
Assistant Professor, Stony Brook University | TTIC, Harvard
Natural Language ProcessingMachine Learning
I
I. V. Ramakrishnan
Stony Brook University
V
Vikas Ashok
Old Dominion University