AI Evaluation Should Measure Verification Cost, Not Correctness Alone
Current AI evaluation frameworks overemphasize output correctness while neglecting the resource costs required to verify errors in real-world deployment, allowing high accuracy metrics to mask substantial verification burdens. This work introduces verification-cost errors (VCEs)—errors that cannot be detected by a specified proportion of validators within a given verification budget—thereby shifting the paradigm from defining errors solely by output properties to centering on their detectability during verification. Through an operational definition, verification budget modeling, and user studies, we empirically demonstrate in code generation and multimodal document understanding tasks that high benchmark accuracy can coexist with significant verification effort, underscoring that correctness alone is insufficient to reflect system reliability in practical settings.