HoliBench: A Cross-Platform Benchmarking and Deployment Toolkit for Foundation Models in CPS-IoT Applications

📅 2026-09-11
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
HoliBench解决跨平台基准测试和部署问题,通过统一的工作流程支持多种模型、推理引擎等,评估准确性、延迟和能耗。
📝 Abstract
Foundation models, including large language models, vision-language models, and time-series foundation models, are increasingly deployed on embedded and edge platforms for CPS and IoT applications, where energy, latency, and memory are as critical as task accuracy. Existing benchmarking tools evaluate model capability in isolation, reporting accuracy assuming sufficient compute, while hardware profiling tools remain platform-specific and mutually incompatible. As a result, users lack a unified workflow for making deployment decisions across heterogeneous devices. We present HoliBench, a modular benchmarking and deployment toolkit that jointly characterizes accuracy, latency, and energy across platforms from single-board computers to GPU servers. Its platform abstraction layer calibrates cross-device measurement, and the toolkit supports multiple model modalities, inference engines, concurrencies, and existing evaluation harnesses. An interactive interface exposes constraint-aware configuration selection over a design space that is profiled once and reused across studies. Using HoliBench, we characterize 20 models across 7 device types, 3 quantization levels, 8 inference backends, and over 30 tasks, surfacing tradeoffs that existing tools miss: quantization reduces latency only on hardware with low-precision support, accuracy gains show diminishing returns relative to energy, and for autoregressive workloads, average inference power is approximately constant across output lengths. We further find that single-model profiles compose under sequential co-resident execution. In a multi-model CPS deployment, standalone profiles predict combined-pipeline latency and power within 1.2% and 2.5%, enabling deployment exploration without exhaustively profiling every pipeline configuration. We release HoliBench as open-source infrastructure for deployment-aware evaluation of foundation models.
Problem

Research questions and friction points this paper is trying to address.

Benchmarking
Deployment
Heterogeneous Devices
Foundation Models
CPS-IoT
Innovation

Methods, ideas, or system contributions that make the work stand out.

Cross-Platform Benchmarking
Deployment Toolkit
Foundation Models
CPS-IoT Applications
Interactive Interface
💼 Related Jobs
No related jobs found.
I
Inesh Chakrabarti
University of California, Los Angeles
Z
Zejun Xiong
University of California, Los Angeles
P
Pragya Sharma
University of California, Los Angeles
Mani Srivastava
Mani Srivastava
Professor of Electrical & Computer Engineering, and Professor of Computer Science, UCLA
Embedded SystemsWireless NetworksCyber-Physical SystemsMobile ComputingSensor Networks