SEATauBench: Adapting Tool-Agent-User Evaluation Into Low-Resource Southeast Asian Languages
This work addresses the inadequacy of existing evaluation frameworks in effectively assessing agent capabilities for low-resource languages in Southeast Asia, which fails to reflect the real-world performance of sovereign AI in local contexts. To bridge this gap, we propose SEATauBench—the first multilingual agent benchmark tailored for Southeast Asian sovereign AI—extending the tool-agent-user evaluation paradigm to this region and introducing a reusable, multi-tiered localization pipeline encompassing conversational language, tool descriptions, and task domains. Cross-lingual transfer experiments based on TauBench reveal that while model performance remains relatively stable when only the conversational language is switched, it degrades substantially as localization depth increases, particularly under full-domain adaptation. These findings underscore the severe limitations of evaluations relying solely on English-centric benchmarks.