🤖 AI Summary
This work proposes Teaching Monster Challenge, the first AI teaching competency benchmark centered on learner adaptability as a core evaluation dimension, designed to assess whether AI agents possess pedagogical content knowledge (PCK) from educational theory. The benchmark requires systems to generate full instructional videos based on a given topic and learner profile, evaluated through a multi-stage framework combining large language model (LLM)-based automated screening, crowdsourced pairwise preference voting, and expert final review. Results indicate that current AI systems perform well in content accuracy but exhibit significant deficiencies in pedagogical delivery and personalized adaptation. While LLM-based scoring effectively identifies low-quality outputs, it struggles to differentiate high-performing systems and shows notable divergence from human preferences in ranking. The study also releases the benchmark dataset, evaluation rubrics, and human assessment results as open resources.
📝 Abstract
AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been benchmarked. To measure it, we introduce the Teaching Monster Challenge, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion. Each system is given a topic and a learner persona and must generate a complete instructional video. Every video is screened by an LLM-judge, ranked by crowd pairwise voting, and finalized by an expert panel. The first edition shows that today's systems handle the content well but are far weaker at presenting it and adapting it to the learner. The same process exposes a limit of automatic judging. The LLM-judge separates a clear low-performing tail but ranks the strongest systems poorly. The strongest systems receive nearly identical scores from the judge, so its ranking of them does not match human preference. Progress therefore requires not only better teaching systems but also better automatic judges, and we release the benchmark, rubric, and human judgments as a testbed for both.