AppEval: A Unified Benchmark for LLM-Based Mobile Application Repair in ArkTS, Swift, and Kotlin

📅 2026-08-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究提出AppEval,针对移动应用修复问题,通过统一基准和原生工具链评估框架,在ArkTS、Swift和Kotlin中测试修复效果,确保修复能跨越构建-安装-启动-测试边界。
📝 Abstract
Repository-level LLM agents are typically evaluated on projects whose tests run on the build host. It remains unclear whether their repairs survive the mobile build-install-launch-test boundary, where a missing SDK, offline device, or pre-assertion crash can be mistaken for a program failure. We present AppEval, a benchmark and native-toolchain evaluation framework for mobile application repair across HarmonyOS/ArkTS, iOS/Swift, and Android/Kotlin. Each task separates a hidden behavior test from the reference production fix and is accepted only when the same installed-app target reaches an assertion failure on the defective revision and passes after the fix; infrastructure failures remain a distinct outcome. A common schema maps this contract to each platform's build system, runtime, and test runner. The audited Android partition contains 200 accepted instrumentation tasks from 24 independently buildable repositories. On these tasks, five agents achieve Pass@1 between 22.00% and 90.50%, a 68.50-percentage-point spread under the same dynamic oracle. These results show that mobile repair performance depends strongly on the evaluated agent while demonstrating why runtime-aware acceptance is necessary for meaningful comparison. The quantitative findings in this paper are Android-specific; audited iOS and HarmonyOS results are required before drawing cross-platform generalization conclusions.
Problem

Research questions and friction points this paper is trying to address.

LLM
mobile application repair
build-install-launch-test boundary
Innovation

Methods, ideas, or system contributions that make the work stand out.

Unified Benchmark
Mobile Application Repair
Runtime-Aware Acceptance
Cross-Platform
🔎 Similar Papers
B
Bang Xie
Shanghai Jiao Tong University, Shanghai, China
Hao Liu
Hao Liu
University of Electronic Science and Technology of China
RISstacked intelligent metasurfaceDRL
Z
Zhenyu Shi
Shanghai Jiao Tong University, Shanghai, China
Yonghao Zhang
Yonghao Zhang
Shanghai Jiao Tong University, Shanghai, China
S
Senjian Zhang
Shanghai Jiao Tong University, Shanghai, China
Zhiyuan Peng
Zhiyuan Peng
Shanghai Jiao Tong University
Code GenerationBlockchain
Xin Yin
Xin Yin
Zhejiang University
LLM4CodeAgent
C
Chenhao Ying
Shanghai Jiao Tong University, Shanghai, China
Y
Yuan Luo
Shanghai Jiao Tong University, Shanghai, China
W
Wei Chen
Shanghai Jiao Tong University, Shanghai, China
Haiming Jin
Haiming Jin
Shanghai Jiao Tong University
wireless sensingreinforcement learning
S
Shaocong Long
Shanghai Jiao Tong University, Shanghai, China
X
Xu Liu
Shanghai Jiao Tong University, Shanghai, China
Zhe Peng
Zhe Peng
Assistant Professor, The Hong Kong Polytechnic University
BlockchainWeb3AIoTData Security and Privacy