Institution profile

Haverford College

Academic institutionnorthamerica · us
Official website
Research library3linked papers
Opportunities0open roles
Selected work

Representative Papers

Benchmarking LLM Competence on Logical Inference over Probability Operators

Jul 29, 2026

This study addresses the difficulty large language models (LLMs) face in distinguishing genuine logical reasoning from superficial pattern matching when processing probabilistic expressions such as “might” or “must,” revealing a fundamental deficiency in their capacity for logical inference under uncertainty. The authors present the first systematic benchmark comprising 14,320 synthetically generated English samples across 15 reasoning templates, carefully controlling variables including question formulation, negation strategies, and surface-level content to isolate models’ logical capabilities with modal expressions. Introducing a “reasoning floor” metric—defined as the worst-case accuracy on Yes/No tasks—the work evaluates whether models exhibit true reasoning rather than response biases. Evaluation of 29 prominent LLMs shows only nine surpass random performance, with most exhibiting significant, logic-irrelevant deviations tied to question phrasing, verb animacy, and named entity attributes.

0 citationsRead paper

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

Jul 25, 2026

This study addresses the limitations of existing cross-lingual reasoning benchmarks, which rely on translated English datasets and thus introduce linguistic bias while failing to assess models’ analogical reasoning within authentic cultural contexts. To overcome this, the authors propose the first language-agnostic, culturally grounded evaluation framework for analogical reasoning. They construct native, high-difficulty analogy datasets for Arabic, Amharic, and Japanese through a collaborative process involving native speakers and large language models, entirely bypassing translation. Experiments across 14 open-source models reveal a stark performance gap: despite strong results on English proverb-based tasks, model accuracy drops by 12–52 percentage points on the localized benchmarks, exposing significant deficiencies in cultural reasoning. The complete pipeline, datasets, and evaluation suite are publicly released.

0 citationsRead paper

Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference

Mar 06, 2025International Conference on Computational Linguistics

This work investigates whether natural language inference (NLI) data generated by large language models (LLMs)—specifically GPT-4, Llama-2-70b, and Mistral-7b—inherit annotation artifacts and societal biases (e.g., gender, race, age) present in human-annotated NLI datasets. We construct LLM-generated NLI subsets and empirically identify severe hypothesis exclusivity bias and stereotypical social biases—first such evidence in synthetic NLI data. To detect these biases systematically, we propose a dual-path framework: (1) a fine-tuned BERT-based hypothesis exclusivity classifier, and (2) pointwise mutual information (PMI) analysis for bias-associated lexical patterns. Experiments show the framework achieves 86–96% classification accuracy on LLM-generated data—significantly outperforming its performance on human-annotated data—and quantitatively identifies multiple bias-correlated lexical terms. Our findings provide both novel methodology and critical empirical evidence for bias assessment and mitigation in LLM-synthesized training data.

0 citationsRead paper
Recent publications

Latest Papers

Benchmarking LLM Competence on Logical Inference over Probability Operators

Jul 29, 2026

This study addresses the difficulty large language models (LLMs) face in distinguishing genuine logical reasoning from superficial pattern matching when processing probabilistic expressions such as “might” or “must,” revealing a fundamental deficiency in their capacity for logical inference under uncertainty. The authors present the first systematic benchmark comprising 14,320 synthetically generated English samples across 15 reasoning templates, carefully controlling variables including question formulation, negation strategies, and surface-level content to isolate models’ logical capabilities with modal expressions. Introducing a “reasoning floor” metric—defined as the worst-case accuracy on Yes/No tasks—the work evaluates whether models exhibit true reasoning rather than response biases. Evaluation of 29 prominent LLMs shows only nine surpass random performance, with most exhibiting significant, logic-irrelevant deviations tied to question phrasing, verb animacy, and named entity attributes.

0 citationsRead paper

ADAGE: A Language-Agnostic Pipeline for Analogical Reasoning Evaluation

Jul 25, 2026

This study addresses the limitations of existing cross-lingual reasoning benchmarks, which rely on translated English datasets and thus introduce linguistic bias while failing to assess models’ analogical reasoning within authentic cultural contexts. To overcome this, the authors propose the first language-agnostic, culturally grounded evaluation framework for analogical reasoning. They construct native, high-difficulty analogy datasets for Arabic, Amharic, and Japanese through a collaborative process involving native speakers and large language models, entirely bypassing translation. Experiments across 14 open-source models reveal a stark performance gap: despite strong results on English proverb-based tasks, model accuracy drops by 12–52 percentage points on the localized benchmarks, exposing significant deficiencies in cultural reasoning. The complete pipeline, datasets, and evaluation suite are publicly released.

0 citationsRead paper

Biases in Large Language Model-Elicited Text: A Case Study in Natural Language Inference

Mar 06, 2025International Conference on Computational Linguistics

This work investigates whether natural language inference (NLI) data generated by large language models (LLMs)—specifically GPT-4, Llama-2-70b, and Mistral-7b—inherit annotation artifacts and societal biases (e.g., gender, race, age) present in human-annotated NLI datasets. We construct LLM-generated NLI subsets and empirically identify severe hypothesis exclusivity bias and stereotypical social biases—first such evidence in synthetic NLI data. To detect these biases systematically, we propose a dual-path framework: (1) a fine-tuned BERT-based hypothesis exclusivity classifier, and (2) pointwise mutual information (PMI) analysis for bias-associated lexical patterns. Experiments show the framework achieves 86–96% classification accuracy on LLM-generated data—significantly outperforming its performance on human-annotated data—and quantitatively identifies multiple bias-correlated lexical terms. Our findings provide both novel methodology and critical empirical evidence for bias assessment and mitigation in LLM-synthesized training data.

0 citationsRead paper