A Study of the Plausibility of Attention between RNN Encoders in Natural Language Inference
Cross-attention maps in natural language inference (NLI) are widely assumed to be interpretable, yet their actual capacity to reveal sentence comparison and logical reasoning processes remains inadequately evaluated—particularly due to the high cost and limited scale of human-annotated explanations. Method: This work introduces, for the first time in NLI, a heuristic rule-based automatic explanation annotation method to overcome these bottlenecks, and conducts a comparative analysis between human and heuristic annotations on the eSNLI dataset. Contribution/Results: Experiments show strong positive correlation (ρ > 0.6) between heuristic and human annotations, validating the former’s utility for explanation quality evaluation. In contrast, raw RNN cross-attention weights exhibit only weak correlation (ρ ≈ 0.2) with human-validated explanations, exposing their limited interpretability. This study establishes a new benchmark, methodology, and empirical insight for assessing attention mechanism interpretability in NLI.