CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty
This work addresses a critical gap in existing benchmarks for large language model (LLM) agents, which often neglect the uncertainty inherent in real-world user interactions and thus fail to assess agent reliability under ambiguous or incomplete requests. To this end, the authors propose the first multi-turn dialogue evaluation benchmark tailored for in-vehicle voice assistants, integrating LLM-simulated users, 58 domain-specific tools, and policy constraints. The benchmark introduces a Disambiguation task to evaluate an agent’s ability to proactively seek clarification and, innovatively, a Hallucination task to probe its awareness of operational boundaries when information is missing. Experimental results reveal that state-of-the-art models achieve less than 50% success rates on the Disambiguation task—often failing due to premature action—and frequently hallucinate or violate policy constraints in the Hallucination task, exposing significant reliability limitations in realistic interactive settings.