Zipbench: Low-Cost Framework for Compressing Comprehensive Benchmarks of Large Language Models

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
为降低大型语言模型评估成本,提出ZipBench框架,通过评估少量基准模型并合成伪评估结果来选择代表性样本,有效压缩了综合基准测试集。
📝 Abstract
Comprehensive benchmark suites are essential for improving large language models (LLMs), but many widely used benchmarks are redundant, making evaluation unnecessarily expensive. Although recent benchmark compression methods (BCMs) can mitigate this cost, many strong BCMs rely on large collections of per-sample evaluation results from numerous LLMs to identify representative samples. Building such collections is also expensive unless they are already public, making these methods difficult to extend to newly released benchmarks. To address this challenge, we present ZipBench, a simple and low-cost BCM with theoretical error and rank-consistency guarantees. ZipBench evaluates only a small set of anchor LLMs, synthesizes pseudo evaluation results to broaden coverage, learns compact sample representations, and selects a small yet representative subset. Building on it, we create ZipBench Zoo, a collection of compact versions of 100+ benchmark proxies spanning text, multimodal, and agent tasks. These benchmark achieve mean absolute errors of 0.002--0.02 and average Spearman correlations of ~0.98 with the full benchmarks. Overall, ZipBench reduces the cost of both LLM evaluation and compact benchmark construction, lowering the barrier to broad LLM research for compute-constrained researchers. The code has been released in https://github.com/MilkThink-Lab/ZipBench.
Problem

Research questions and friction points this paper is trying to address.

large language models
benchmark compression
evaluation cost
Innovation

Methods, ideas, or system contributions that make the work stand out.

Benchmark Compression
Low-Cost Evaluation
Pseudo Evaluation Results
Compact Sample Representations
Rank-Consistency Guarantees
🔎 Similar Papers
No similar papers found.
💼 Related Jobs
No related jobs found.
Z
Zhongzhan Huang
Bosch Research
J
Junxin Li
Bosch Research, Sun Yat-sen University
G
Guoming Ling
Sun Yat-sen University
Y
Yupei Lin
Sun Yat-sen University
Shanshan Zhong
Shanshan Zhong
Carnegie Mellon University
Language ModelsMultimodal UnderstandingMultimodal Generation
Hefeng Wu
Hefeng Wu
Sun Yat-sen University
Computer visionMachine LearningArtificial Intelligence