EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文通过EduFair-Bench评估LLM教师在不同学生群体中的教学公平性,使用多维度模拟和统计测试方法检测偏见。
📝 Abstract
Large language models (LLMs) are increasingly deployed as tutors, but it is unclear whether they support all students equally well. We introduce \textbf{EduFair-Bench}, a benchmark for auditing the pedagogical fairness of LLM tutors---whether tutoring quality varies systematically with student demographics. EduFair-Bench pairs a multi-domain question bank (mathematics, physics, chemistry) with a controlled simulation in which a fixed LLM student interacts with each tutor across nine demographic levels spanning four dimensions: gender, immigration background, first language, and socioeconomic status (SES). Tutoring quality is scored on five turn-level pedagogical metrics and four conversation-level dimensions, using an LLM judge validated against three-annotator consensus on 180 tutor turns. Bias is measured via paired Wilcoxon signed-rank tests and bootstrap effect-size confidence intervals. Two ablations (demographic cues conveyed through names; conflicting demographic information between tutor and student) disentangle tutor-driven from student-driven bias. Across five tutors, we find that model capability and demographic fairness are largely orthogonal: the smallest model is the most consistent while the four more capable tutors all exhibit wide demographic gaps with no clear capability-to-fairness ordering, pedagogy-specific RL training redistributes rather than removes bias, and language- and immigration-related cues produce larger gaps than gender- and SES-related cues.
Problem

Research questions and friction points this paper is trying to address.

EduFair-Bench
pedagogical fairness
student demographics
large language models
Innovation

Methods, ideas, or system contributions that make the work stand out.

EduFair-Bench
pedagogical fairness
demographic bias
LLM tutors
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Jiaxu Zhao
Jiaxu Zhao
Eindhoven University of Technology
Large Language ModelsAI SafetyCausal Reasoning
B
Bahar Radmehr
Swiss Federal Institute of Technology in Lausanne (EPFL), Switzerland
F
Fares Fawzi
Swiss Federal Institute of Technology in Lausanne (EPFL), Switzerland
T
Tanya Nazaretsky
Swiss Federal Institute of Technology in Lausanne (EPFL), Switzerland
Tanja Käser
Tanja Käser
Tenure Track Assistant Professor
Educational Data MiningAI for EducationLearning Analytics