CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows

📅 2026-09-07
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
CausalVerify通过执行验证解决LLM在因果推断工作流中估计准确性的问题,使用合成数据集和实验评估方法的有效性。
📝 Abstract
Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional context) with 100 fixed-seed synthetic scenarios that realise CSV datasets for difference-in-differences, event study, instrumental variables, and regression discontinuity designs. Experiment A (real-paper text agreement) scores method-family and direction agreement against four-LLM consensus labels. Experiment B (synthetic execution) runs model-written R code and checks whether the extracted treatment-effect estimate matches a canonical estimator on the same realised dataset; this execution-grounded correctness layer is L2b+, distinct from L2b, which records only whether the code executes. A calibration arm asks whether self-reported confidence separates correct from incorrect workflows. On Experiment B, seven LLMs reach L2b+ pass rates of 10% to 88% at the default 50% tolerance, and 66 of the 426 workflows that execute (15.5%) return a wrong estimate. Execution ranking (L2b) agrees with L2b+ far better than text-direction scoring (L4): Kendall $\tau=0.81$ and Spearman $\rho=0.93$, versus Kendall $\tau$ between $-0.20$ and $0.10$ for L4. Llama-3.3-70B-Instruct shows the same qualitative gap, and reported confidence does not reliably separate correct from incorrect workflows. The claims are confined to standardized single-shot workflows in these four design families under the evaluated R backend and model panel; the benchmark does not measure general causal-inference ability. Code, data, cached outputs, and a datasheet are released.
Problem

Research questions and friction points this paper is trying to address.

causal inference
large language models
benchmark
execution verification
econometrics
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Inference
Benchmarking
Execution-Grounded Verification
Synthetic Scenarios
Econometric Workflows
Y
Yonghong Zhang
Department of Finance, Universidad Autónoma de Madrid, Madrid, Spain
Ricardo Correia
Ricardo Correia
Department of Finance, Universidad Autónoma de Madrid, Madrid, Spain
I
Isabel M. Parra
Department of Finance, Universidad Autónoma de Madrid, Madrid, Spain
Y
Yong Xie
Spanish National Research Council (CSIC), Madrid, Spain