A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance

📅 2026-09-14
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
本文提出一种结合专家意见与合成数据生成的方法,以解决现有大规模语言模型评估基准在有效性和可扩展性之间的权衡问题。
📝 Abstract
This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation. Existing benchmark construction methods often trade off validity and scalability: datasets designed with domain experts can produce high-quality evaluations but are slow and costly to create, while synthetically generating data may scale efficiently but often results in unrealistic, redundant, or out-of-scope examples. To address this gap, we introduce a schema eliciting key information about the goals, scope, and context of an evaluation task, and use this information to guide synthetic data generation. We further define four criteria grounded in measurement validity for assessing dataset quality: coverage, diversity, content realism, and stylistic realism. Using these criteria, we show how expert-informed scaffolds can guide synthetic data generation toward more valid benchmarks. Through quantitative evaluations and a real-world case study with domain experts, we demonstrate that our approach improves benchmark data quality over existing methods while preserving validity. We additionally analyze how different types of schema information affect different dataset quality criteria, and provide practical guidance on which information to prioritize collecting under resource constraints.
Problem

Research questions and friction points this paper is trying to address.

context-specific benchmarks
expert guidance
synthetic data generation
validity and scalability
Innovation

Methods, ideas, or system contributions that make the work stand out.

end-to-end approach
context-specific benchmarks
synthetic data generation
measurement validity
🔎 Similar Papers
No similar papers found.