GGUF-Metadata Prediction of Single-Sequence llama.cpp Throughput Across Three Systems

📅 2026-09-13
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
研究通过使用特定量化因子的roofline形预测器,基于GGUF元数据预测单序列模型吞吐量,解决了不同系统上模型性能预测问题。
📝 Abstract
We predict single-sequence model throughput from GGUF metadata using roofline-shaped predictors with quantization-specific scale factors fitted on reference models. The scored cohort comprises 318 phase-depth measurements from 53 host-file configurations on two Apple M4 Max systems and an NVIDIA RTX 5080. On host-specific held-out sets of four, five, and two configurations, an active-parameter decode model obtains 13.1%, 14.4%, and 36.1% mean absolute percentage error (MAPE), versus 49.4%, 55.3%, and 51.9% when charging total parameters. Leave-one-host-out coefficients fitted on the other two systems yield 11.6%, 16.8%, and 36.0% test MAPE. A low-bit model ladder changes ordering across runtime stacks. The P2 prefill baseline gives 18.7%, 22.2%, and 108.2% test MAPE. GGUF structure helps on all three systems, but fitted efficiencies are not universal.
Problem

Research questions and friction points this paper is trying to address.

throughput
GGUF metadata
single-sequence model
quantization-specific scale factors
roofline-shaped predictors
Innovation

Methods, ideas, or system contributions that make the work stand out.

GGUF metadata
roofline-shaped predictors
quantization-specific scale factors
🔎 Similar Papers
No similar papers found.
X
Xinyu Qiu
Northeastern University
C
Chuhong Xu
Sofia University
B
Bo Su
Indiana University
Ziyao Chen
Ziyao Chen
University of California, San Diego
R
Ruiyang Xu
Northeastern University
S
Shimeng Dai
Michigan State University