From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers

📅 2026-08-24
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究评估了12种语言模型在因果关系判断上的可靠性,发现这些模型倾向于过度预测因果边,并且对直接因果关系的识别不够准确。传统的置信度估计不可靠,而跨提示和跨模型的一致性则提供了更好的信号。
📝 Abstract
Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: verbalized, logit-based, cross-prompt agreement, and cross-model agreement. Under our language-only pairwise protocol, our evaluation yields three key findings. (i) LLM-based causal judgments are strongly recall-dominant: models predict overly dense graphs with many false-positive edges, while prompting mainly shifts the precision-recall trade-off rather than resolving overprediction. Gains from model scale diminish on the largest graphs and do not eliminate miscalibration. (ii) LLMs often capture causal relatedness without reliably identifying directness or orientation. Relative to published reference graphs, models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges, versus 28.2% of other non-edges. Moreover, 80.8% and 84.6% of these false positives receive verbalized confidence of at least 80%, revealing substantial overconfidence in structurally incorrect predictions. (iii) Conventional confidence estimates are unreliable, whereas agreement offers a more promising signal. Logit-based confidence frequently collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement achieve better mean calibration and discrimination, though their advantages are not statistically significant after Holm correction. A benchmark-familiarity audit further identifies potential familiarity in five model-dataset pairs, all involving AsiaM. Overall, our results suggest LLMs are better viewed as sources of externally validated soft causal priors than as direct evidence of causal structure.
Problem

Research questions and friction points this paper is trying to address.

Causal Judgments
Large Language Models
Confidence Calibration
Structural Causal Discovery
False Positives
Innovation

Methods, ideas, or system contributions that make the work stand out.

Causal Edge Classification
Model Calibration
Cross-Model Agreement
Prompting Strategies
🔎 Similar Papers
A
Amit Kumar
Texas A&M University-Corpus Christi, USA
E
Elnur Adl Zarabi
Texas A&M University-Corpus Christi, USA
S
Suranjana Trivedy
BITS Pilani Goa, India
Z
Zhiqian Chen
Mississippi State University, USA
L
Lei Zhang
Northern Illinois University, USA
Kaiqun Fu
Kaiqun Fu
Texas Christian University
spatial data mininggraph neural networksmachine learningurban computing
Taoran Ji
Taoran Ji
Assistant Professor - Texas A&M University - Corpus Christi | Ex - Moody's Analytics
Machine LearningTime-Series ForecastingNatural Language ProcessingEvent DetectionEvent