Institution profile

Mercor

Industry researchnorthamerica · us
Official website
Research library6linked papers
Opportunities0open roles
Selected work

Representative Papers

APEX-Accounting

Jul 29, 2026

This study evaluates the practical capabilities of state-of-the-art large language models on real-world accounting tasks, including reconciliation, accruals, journal entry preparation, and financial statement generation. To this end, we introduce the first high-quality, closed-domain benchmark tailored to accounting practice, comprising 10 synthetic enterprises and 160 expert-designed and scored tasks that support multimodal financial documents such as PDFs and spreadsheets, all evaluated under a fixed token budget. Our analysis uncovers a Simpson’s paradox between model performance and token consumption. While Claude-Fable-5 (Max) achieves the highest Mean Criteria@3 score at 56.4%, no model exceeds 21.5% on Pass@8, revealing substantial limitations in current models’ ability to perform complex accounting reasoning.

0 citationsRead paper

APEX-Agents

Jan 20, 2026

This study evaluates the capability of AI agents to perform long-horizon, cross-application complex tasks in professional service domains such as investment banking, management consulting, and corporate law. To this end, we introduce Archipelago, the first high-fidelity benchmark tailored to the professional services industry, comprising 480 real-world task scenarios derived from authentic workflows, along with an automated execution framework and human-defined scoring rubrics. Using the Pass@1 metric, we benchmark leading AI agents and find that Gemini 3 Flash (with Thinking=High) achieves the highest performance at 24.0%. The complete dataset and evaluation infrastructure are publicly released to establish a new standard for agent research in specialized professional domains.

0 citationsRead paper

APEX-SWE

Jan 13, 2026

Current AI evaluation methodologies are largely confined to narrow tasks and fail to assess models’ capacity to perform high-value work in real-world software engineering contexts. This work proposes the APEX-SWE benchmark, establishing the first evaluation paradigm centered on authentic software engineering workflows. It evaluates models through two challenge types—integrated tasks and observability tasks—to probe cognitive reasoning and proactive decision-making in complex, open-ended environments. The assessment incorporates end-to-end system integration, cloud-native interactions, and telemetry signal analysis, combining both structured and unstructured contextual information. Among eight state-of-the-art models evaluated, Gemini 3 Pro (with Thinking=High) achieves the highest performance (Pass@1 of 25%), owing to its ability to effectively distinguish hypotheses from facts and actively resolve uncertainty.

0 citationsRead paper

The AI Consumer Index (ACE)

Dec 04, 2025

This work addresses the limited practical utility of state-of-the-art AI models in high-stakes consumer tasks—shopping, dining, gaming, and DIY. To bridge this gap, we introduce ACE, the first dedicated benchmark for evaluating consumer-oriented AI capabilities. We propose dynamic provenance verification, an automated method to assess whether model responses faithfully reflect retrieved web content, and integrate a hybrid human–automated, tiered evaluation framework that rigorously tests the factual accuracy and operational viability of critical elements (e.g., prices, URLs). We release a hidden test set of 400 instances and an open-source development set, filling a critical void in consumer-grade AI evaluation. Comprehensive evaluation of 10 SOTA models reveals a maximum overall score of only 56.1%; even the best-performing shopping model scores below 50%. Widespread hallucination and non-executable outputs confirm that current AI systems remain substantially inadequate for real-world consumer applications.

0 citationsRead paper
Recent publications

Latest Papers

APEX-Accounting

Jul 29, 2026

This study evaluates the practical capabilities of state-of-the-art large language models on real-world accounting tasks, including reconciliation, accruals, journal entry preparation, and financial statement generation. To this end, we introduce the first high-quality, closed-domain benchmark tailored to accounting practice, comprising 10 synthetic enterprises and 160 expert-designed and scored tasks that support multimodal financial documents such as PDFs and spreadsheets, all evaluated under a fixed token budget. Our analysis uncovers a Simpson’s paradox between model performance and token consumption. While Claude-Fable-5 (Max) achieves the highest Mean Criteria@3 score at 56.4%, no model exceeds 21.5% on Pass@8, revealing substantial limitations in current models’ ability to perform complex accounting reasoning.

0 citationsRead paper

APEX-Agents

Jan 20, 2026

This study evaluates the capability of AI agents to perform long-horizon, cross-application complex tasks in professional service domains such as investment banking, management consulting, and corporate law. To this end, we introduce Archipelago, the first high-fidelity benchmark tailored to the professional services industry, comprising 480 real-world task scenarios derived from authentic workflows, along with an automated execution framework and human-defined scoring rubrics. Using the Pass@1 metric, we benchmark leading AI agents and find that Gemini 3 Flash (with Thinking=High) achieves the highest performance at 24.0%. The complete dataset and evaluation infrastructure are publicly released to establish a new standard for agent research in specialized professional domains.

0 citationsRead paper

APEX-SWE

Jan 13, 2026

Current AI evaluation methodologies are largely confined to narrow tasks and fail to assess models’ capacity to perform high-value work in real-world software engineering contexts. This work proposes the APEX-SWE benchmark, establishing the first evaluation paradigm centered on authentic software engineering workflows. It evaluates models through two challenge types—integrated tasks and observability tasks—to probe cognitive reasoning and proactive decision-making in complex, open-ended environments. The assessment incorporates end-to-end system integration, cloud-native interactions, and telemetry signal analysis, combining both structured and unstructured contextual information. Among eight state-of-the-art models evaluated, Gemini 3 Pro (with Thinking=High) achieves the highest performance (Pass@1 of 25%), owing to its ability to effectively distinguish hypotheses from facts and actively resolve uncertainty.

0 citationsRead paper

The AI Consumer Index (ACE)

Dec 04, 2025

This work addresses the limited practical utility of state-of-the-art AI models in high-stakes consumer tasks—shopping, dining, gaming, and DIY. To bridge this gap, we introduce ACE, the first dedicated benchmark for evaluating consumer-oriented AI capabilities. We propose dynamic provenance verification, an automated method to assess whether model responses faithfully reflect retrieved web content, and integrate a hybrid human–automated, tiered evaluation framework that rigorously tests the factual accuracy and operational viability of critical elements (e.g., prices, URLs). We release a hidden test set of 400 instances and an open-source development set, filling a critical void in consumer-grade AI evaluation. Comprehensive evaluation of 10 SOTA models reveals a maximum overall score of only 56.1%; even the best-performing shopping model scores below 50%. Widespread hallucination and non-executable outputs confirm that current AI systems remain substantially inadequate for real-world consumer applications.

0 citationsRead paper