Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents

📅 2026-08-09
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work proposes Teaching Monster Challenge, the first AI teaching competency benchmark centered on learner adaptability as a core evaluation dimension, designed to assess whether AI agents possess pedagogical content knowledge (PCK) from educational theory. The benchmark requires systems to generate full instructional videos based on a given topic and learner profile, evaluated through a multi-stage framework combining large language model (LLM)-based automated screening, crowdsourced pairwise preference voting, and expert final review. Results indicate that current AI systems perform well in content accuracy but exhibit significant deficiencies in pedagogical delivery and personalized adaptation. While LLM-based scoring effectively identifies low-quality outputs, it struggles to differentiate high-performing systems and shows notable divergence from human preferences in ranking. The study also releases the benchmark dataset, evaluation rubrics, and human assessment results as open resources.
📝 Abstract
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been benchmarked. To measure it, we introduce the Teaching Monster Challenge, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion. Each system is given a topic and a learner persona and must generate a complete instructional video. Every video is screened by an LLM-judge, ranked by crowd pairwise voting, and finalized by an expert panel. The first edition shows that today's systems handle the content well but are far weaker at presenting it and adapting it to the learner. The same process exposes a limit of automatic judging. The LLM-judge separates a clear low-performing tail but ranks the strongest systems poorly. The strongest systems receive nearly identical scores from the judge, so its ranking of them does not match human preference. Progress therefore requires not only better teaching systems but also better automatic judges, and we release the benchmark, rubric, and human judgments as a testbed for both.
Problem

Research questions and friction points this paper is trying to address.

Pedagogical Content Knowledge
AI agents
instructional video generation
learner adaptation
benchmarking
Innovation

Methods, ideas, or system contributions that make the work stand out.

Pedagogical Content Knowledge
Instructional Video Generation
Learner Persona Adaptation
AI Teaching Benchmark
LLM-based Evaluation
Yi-Cheng Lin
Yi-Cheng Lin
National Taiwan University
Speech ProcessingMachine LearningFairness
Y
Yu-Kai Guo
National Taiwan University
S
Szu-Chi Chen
National Taiwan University
B
Bo-Han Feng
National Taiwan University
Y
Yun-Man Hsu
National Taiwan University
H
Hsiang Hsieh
National Taiwan University
Y
Yu-Jung Lin
National Taiwan University
Y
Yue-Ling Wu
National Taiwan University
J
Jia-Kai Dong
National Taiwan University
A
An-Yu Cheng
National Taiwan University
Y
Yu-Han Huang
National Taiwan University
L
Lok-Lam Ieong
National Taiwan University
Kuan-Yu Chen
Kuan-Yu Chen
National Taiwan University of Science and Technology
Language ModelingSpeech RecognitionInformation RetrievalSummarizationNature Language Processing
M
Ming-Douo Tchouang
National Taiwan University
Shao-Hua Sun
Shao-Hua Sun
Assistant Professor at National Taiwan University
Machine LearningRobot LearningReinforcement LearningProgram Synthesis
Che Lin
Che Lin
National Taiwan University
Deep learningData scienceMedical AIFinTechSignal processing in wireless communications.
J
Jian-Jiun Ding
National Taiwan University
Hung-yi Lee
Hung-yi Lee
National Taiwan University
deep learningspoken language understandingspeech processing