The Box is in the Pen: Evaluating Commonsense Reasoning in Neural Machine Translation
This work investigates neural machine translation (NMT) models’ capacity to resolve lexical and syntactic ambiguities via commonsense reasoning. To this end, we introduce CReM—the first commonsense reasoning evaluation benchmark tailored for NMT—comprising 1,200 triplets spanning seven categories of commonsense knowledge. We propose a dual-dimensional evaluation framework assessing both accuracy (where mainstream models achieve only 60.1%) and cross-context consistency (with merely 31% consistency), enabling the first systematic quantification of NMT’s commonsense reasoning capability. Extensive comparative experiments are conducted using models including BERT and GPT-2; statistical analysis and error attribution reveal that contextual modeling and effective commonsense integration remain critical bottlenecks. The CReM benchmark is publicly released to serve as a standardized evaluation tool for future research.