Agentic Confidence Calibration
Existing AI agents often fail in complex tasks due to overconfidence, and static calibration methods struggle to address compounded errors, tool uncertainty, and opaque failures along execution trajectories. This work formalizes the agent confidence calibration problem for the first time and introduces the Holistic Trajectory Calibration (HTC) framework, which models both macro-level dynamics and micro-level stability of execution trajectories to enable interpretable, cross-domain, and generalizable process-level calibration. The proposed General Agent Calibrator (GAC) consistently outperforms strong baselines across eight benchmarks, diverse large language models, and agent architectures, achieving state-of-the-art performance in Expected Calibration Error (ECE) and discrimination ability—particularly excelling on the cross-domain GAIA benchmark.