Adaptation Fidelity of SPEC CPU2026
本文通过对比SPEC CPU2026基准测试与其上游开源版本在单副本和多副本运行场景下的性能,量化分析了两者之间的'保真度差距',验证了SPEC方法的有效性。
本文通过对比SPEC CPU2026基准测试与其上游开源版本在单副本和多副本运行场景下的性能,量化分析了两者之间的'保真度差距',验证了SPEC方法的有效性。
本文介绍了AmpereOne CPU核心性能验证方法,通过周期精确的相关性检查、数据驱动的工作负载管理和高频回归系统等手段,确保处理器达到性能目标。
This study addresses the significant dependence of provenance-based intrusion detection system (PIDS) evaluations on dataset and protocol choices, which often leads to misleading performance comparisons. Conducting a systematic re-evaluation of representative PIDS under a unified temporal split testing protocol and hyperparameter tuning restricted to the validation set—using publicly available datasets that satisfy auditability, labeling, and calibration requirements—the authors find that most reported performance gains stem from lexical novelty in executable names or paths rather than sophisticated provenance modeling. They propose quantifying dataset semantic signal quality via feature completeness and field entropy, which explain model sensitivity to architectural choices. On three of four widely used datasets, a simple allowlist matches or outperforms learning-based methods; only Theia, exhibiting the strongest semantic signals, effectively reveals model advantages in alert prioritization and node recovery.
This work addresses the limitations of existing CPU benchmarks in accurately evaluating the performance of modern heterogeneous, multithreaded processors under diverse workloads. To this end, the authors present the SPEC CPU 2026 benchmark suite, developed through community collaboration and principled methodology, which introduces the Rolling-Round-Robin Rate approach to standardize the execution of heterogeneous multiprogrammed workloads. The suite incorporates newly designed multithreaded benchmarks exhibiting varied microarchitectural characteristics, selected and hardened through an open-source application curation process. Emphasizing workload diversity, portability, and long-term viability, SPEC CPU 2026 establishes a robust, representative, and authoritative standard for performance evaluation, thereby supporting next-generation computer architecture research.
This study addresses the lack of systematic evaluation of ARM Memory Tagging Extension (MTE) performance under real hardware and diverse workloads, particularly across different microarchitectures and security use cases. For the first time, we comprehensively quantify MTE’s performance overhead on multiple ARM processors—including Google Pixel 8/9, AmpereOne, and Apple M5—using a mix of general-purpose and server workloads such as SPEC CPU, RocksDB, and Nginx. We further analyze MTE’s applicability in memory safety, sandboxing, and control-flow integrity mechanisms. Through real-platform profiling and microarchitectural bottleneck analysis, we identify the root causes of slowdowns up to 6.64×, correct methodological flaws in prior studies, and delineate MTE’s practical boundaries: overhead is generally manageable in most scenarios, certain workloads demand hardware-level optimizations, and several security applications are already deployment-ready.
本文通过对比SPEC CPU2026基准测试与其上游开源版本在单副本和多副本运行场景下的性能,量化分析了两者之间的'保真度差距',验证了SPEC方法的有效性。
本文介绍了AmpereOne CPU核心性能验证方法,通过周期精确的相关性检查、数据驱动的工作负载管理和高频回归系统等手段,确保处理器达到性能目标。
This study addresses the significant dependence of provenance-based intrusion detection system (PIDS) evaluations on dataset and protocol choices, which often leads to misleading performance comparisons. Conducting a systematic re-evaluation of representative PIDS under a unified temporal split testing protocol and hyperparameter tuning restricted to the validation set—using publicly available datasets that satisfy auditability, labeling, and calibration requirements—the authors find that most reported performance gains stem from lexical novelty in executable names or paths rather than sophisticated provenance modeling. They propose quantifying dataset semantic signal quality via feature completeness and field entropy, which explain model sensitivity to architectural choices. On three of four widely used datasets, a simple allowlist matches or outperforms learning-based methods; only Theia, exhibiting the strongest semantic signals, effectively reveals model advantages in alert prioritization and node recovery.
This work addresses the limitations of existing CPU benchmarks in accurately evaluating the performance of modern heterogeneous, multithreaded processors under diverse workloads. To this end, the authors present the SPEC CPU 2026 benchmark suite, developed through community collaboration and principled methodology, which introduces the Rolling-Round-Robin Rate approach to standardize the execution of heterogeneous multiprogrammed workloads. The suite incorporates newly designed multithreaded benchmarks exhibiting varied microarchitectural characteristics, selected and hardened through an open-source application curation process. Emphasizing workload diversity, portability, and long-term viability, SPEC CPU 2026 establishes a robust, representative, and authoritative standard for performance evaluation, thereby supporting next-generation computer architecture research.
This study addresses the lack of systematic evaluation of ARM Memory Tagging Extension (MTE) performance under real hardware and diverse workloads, particularly across different microarchitectures and security use cases. For the first time, we comprehensively quantify MTE’s performance overhead on multiple ARM processors—including Google Pixel 8/9, AmpereOne, and Apple M5—using a mix of general-purpose and server workloads such as SPEC CPU, RocksDB, and Nginx. We further analyze MTE’s applicability in memory safety, sandboxing, and control-flow integrity mechanisms. Through real-platform profiling and microarchitectural bottleneck analysis, we identify the root causes of slowdowns up to 6.64×, correct methodological flaws in prior studies, and delineate MTE’s practical boundaries: overhead is generally manageable in most scenarios, certain workloads demand hardware-level optimizations, and several security applications are already deployment-ready.