🤖 AI Summary
研究使用低成本硬件对NEAR协议Nightshade分片进行实证分析,通过改变分片数量来测量性能指标,揭示了不同瓶颈机制。
📝 Abstract
NEAR Protocol's Nightshade architecture targets one million transactions per second (TPS) through horizontal sharding of both state and computation. Published benchmarks were produced on expensive Google Cloud Platform infrastructure costing approximately \$700 per hour, leaving a significant reproducibility gap for academic research. We present the first independent empirical characterization of NEAR Nightshade sharding on commodity hardware: a Chameleon Cloud bare-metal node with 48 hyperthreaded Intel Xeon cores, 128\,GB RAM, and HDD storage at 80--100\,MB/s. We systematically sweep shard count from $N{=}1$ to $N{=}24$, measuring aggregate TPS, per-shard TPS, block time, BFT finality, memory, and disk I/O. We identify three distinct bottleneck regimes: L3 cache pressure at low $N$, witness gossip pipeline saturation at mid $N$, and coherence collapse at high $N$. A key unexpected finding is that HDD write latency acts as implicit flow control for the witness gossip pipeline. Removing it via RAM-backed tmpfs causes complete chain stall at $N{=}16$, with a 29$\times$ spike in orphan witness rate at 47\% CPU utilization. Aggregate TPS peaks at $N{=}8$ (+40\% over $N{=}1$) then reverses, with per-shard TPS collapsing 23$\times$ by $N{=}24$. Our dataset provides the first commodity-hardware calibration baseline for the companion SimPy sharding simulator.