ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

📅 2026-08-10
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
Current evaluations of educational large language models lack comprehensive assessment across multiple critical dimensions, including factual accuracy, safety, pedagogical utility, and alignment with educational objectives. This work proposes ELBench, the first unified benchmark integrating four key dimensions: general capabilities, safety and trustworthiness, foundational education, and higher-order competencies. Combining both public and synthetically generated data, ELBench employs structured task design, modular scoring, and cross-model comparison to systematically evaluate nine mainstream models. The study reveals that Chinese models lead in safety performance, yet domain-specific educational models do not significantly outperform general-purpose counterparts. All models exhibit systematic deficiencies in higher-order competency tasks, and a negative correlation emerges between safety and pedagogical utility, highlighting critical bottlenecks in the current development of educational large language models.
📝 Abstract
Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does not: a usable education-facing model is supposed to be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation, so none assesses education-facing suitability as an integrated profile. We introduce ELBench, the first benchmark to evaluate all four requirements (General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation) on the same models under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. We evaluate nine models, seven frontier general-purpose systems and two education-specialized variants, and report three findings. First, module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable on overall score, yet their module leaders differ substantially, and safety is anti-correlated with practical teaching (r = -0.83). Second, the Chinese-developed models lead the safety module, the most discriminative in the suite; this advantage is largest on region-specific normative content and narrows, but does not vanish, on universal-harm content. Third, the two education-specialized models lead neither education module, and on High-Level Cultivation all models share a systematic blind spot: on the structured judgment task they converge on the same non-reference option, favoring pedagogical style over fit to the stated goal, so the module scores uniformly low and does not separate models. This raises, but does not resolve, whether domain post-training keeps pace with frontier systems on education tasks.
Problem

Research questions and friction points this paper is trying to address.

education-facing LLMs
multi-dimensional benchmark
safety and trustworthiness
pedagogical alignment
model evaluation
Innovation

Methods, ideas, or system contributions that make the work stand out.

multi-dimensional benchmark
education-facing LLMs
safety and pedagogy alignment
high-level cultivation
integrated evaluation framework
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Y
Yilin Jiang
East China Normal University; The Hong Kong University of Science and Technology (Guangzhou)
X
Xiaorong Zhu
Shanghai Jiao Tong University; Shanghai Artificial Intelligence Laboratory
Fei Tan
Fei Tan
Associate Professor, East China Normal University
NLPData MiningNetwork Science
Zicheng Zhang
Zicheng Zhang
Shanghai AI Lab
Multi-modal LLMQuality assessment
K
Kaiyi Huang
East China Normal University
Y
Yang Yu
East China Normal University
Z
Zexuan Fei
East China Normal University
Yiming Luo
Yiming Luo
PhD student, The University of Hong Kong
Robotics
Keqian Li
Keqian Li
GenAI, Meta
Data miningmachine learning
Hao Hao
Hao Hao
East China Normal University
A
Aimin Zhou
East China Normal University
Guangtao Zhai
Guangtao Zhai
Professor, IEEE Fellow, Shanghai Jiao Tong University
Multimedia Signal ProcessingVisual Quality AssessmentQoEAI EvaluationDisplays