Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework
This work addresses the deployment blind spots commonly encountered in developing production-grade AI customer service agents for hundreds of millions of users, which often stem from fragmented workflows across evaluation, context engineering, training, and online metrics. To bridge this gap, the authors propose the first unified development framework centered on evaluation-driven design. The framework integrates structured context engineering, human-in-the-loop prompt iteration, an LLM-based judging mechanism with consistency guarantees (including GEPA optimization), and an end-to-end validation pipeline, achieving strong alignment between offline metrics and online performance. Evaluated across five real-world customer service scenarios, the approach substantially enhances user experience—evidenced by a 37-percentage-point increase in AI Net Promoter Score and a 29-percentage-point rise in self-service rate in the card delivery scenario—with most scenarios approaching expert human-level performance.