Institution profile

DataGrand Information Technology

Industry researchasia · cn
Official website
Research library1linked papers
Opportunities0open roles
Selected work

Representative Papers

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches

May 18, 2025

Existing LLM evaluation suffers from insufficient fairness, poor scalability, and data contamination. To address these challenges, we propose Teach2Eval—a novel indirect evaluation framework grounded in pedagogical capability: it prompts LLMs to act as “teachers” instructing weaker student models on target tasks, then automatically converts their teaching outputs into standardized multiple-choice questions (MCQs). This enables dynamic, contamination-resistant, and scalable automated assessment. Teach2Eval pioneers a cognitive ability measurement paradigm wherein teaching efficacy serves as a proxy metric—overcoming limitations of static benchmarks while ensuring fairness, interpretability, and orthogonality across cognitive dimensions. Evaluated on 26 mainstream LLMs, Teach2Eval achieves strong rank correlation with human judgments and model-level dynamic rankings. Moreover, it supports fine-grained training feedback, substantially enhancing the guidance value and interpretability of LLM evaluation.

0 citationsRead paper
Recent publications

Latest Papers

Teach2Eval: An Indirect Evaluation Method for LLM by Judging How It Teaches

May 18, 2025

Existing LLM evaluation suffers from insufficient fairness, poor scalability, and data contamination. To address these challenges, we propose Teach2Eval—a novel indirect evaluation framework grounded in pedagogical capability: it prompts LLMs to act as “teachers” instructing weaker student models on target tasks, then automatically converts their teaching outputs into standardized multiple-choice questions (MCQs). This enables dynamic, contamination-resistant, and scalable automated assessment. Teach2Eval pioneers a cognitive ability measurement paradigm wherein teaching efficacy serves as a proxy metric—overcoming limitations of static benchmarks while ensuring fairness, interpretability, and orthogonality across cognitive dimensions. Evaluated on 26 mainstream LLMs, Teach2Eval achieves strong rank correlation with human judgments and model-level dynamic rankings. Moreover, it supports fine-grained training feedback, substantially enhancing the guidance value and interpretability of LLM evaluation.

0 citationsRead paper