APEX-Accounting
This study evaluates the practical capabilities of state-of-the-art large language models on real-world accounting tasks, including reconciliation, accruals, journal entry preparation, and financial statement generation. To this end, we introduce the first high-quality, closed-domain benchmark tailored to accounting practice, comprising 10 synthetic enterprises and 160 expert-designed and scored tasks that support multimodal financial documents such as PDFs and spreadsheets, all evaluated under a fixed token budget. Our analysis uncovers a Simpson’s paradox between model performance and token consumption. While Claude-Fable-5 (Max) achieves the highest Mean Criteria@3 score at 56.4%, no model exceeds 21.5% on Pass@8, revealing substantial limitations in current models’ ability to perform complex accounting reasoning.