When Is the Same Model Not the Same Service? A Measurement Study of Hosted Open-Weight LLM APIs

📅 2026-05-04

📈 Citations: 0

✨ Influential: 0

career value

227K/year

🤖 AI Summary

This study addresses the significant gap between the static capabilities of open-source large language models and their real-world performance in hosted API services, a discrepancy exacerbated by insufficient understanding of service-layer heterogeneity and dynamics. Leveraging multidimensional data collected in Q4 2025 from the AI Ping platform—including request logs, metadata, compatibility probes, price snapshots, and latency measurements—the work uncovers three key patterns: demand concentration inertia, supply-demand misalignment, and task-conditioned routing. Building on these insights, the paper reframes model deployment as a constrained statistical decision problem and demonstrates that intelligent routing reduces inference costs for Qwen3-32B by 37.8% and increases throughput for DeepSeek-V3.2 by approximately 90%.

📝 Abstract

Open-weight large language models (LLMs) are often described as downloadable model artifacts, but in production they are increasingly consumed as hosted APIs. This paper studies the intermediary service layer that turns a model release into an operational endpoint. Using sampled request logs, provider metadata, compatibility probes, pricing snapshots, and continuous latency measurements collected by AI Ping during Q4 2025, we analyze demand concentration, provider heterogeneity, and task-conditioned routing for popular open-weight model families. The first empirical pattern is concentration with inertia: among the model families displayed in the public aggregate, the largest family carries 32.0% of relative demand and the top five carry 87.4%, with a Gini coefficient of 0.693, yet older versions remain active after newer releases. The second pattern is a separation between supply and use: broad provider listing of a model does not imply realized adoption, and listed prices are more anchored than latency, throughput, context length, protocol support, and error semantics. The third pattern is conditionality: applications induce different token-length regimes, so the relevant service object is not a model name but a provider-model-task-time tuple under protocol and context constraints. In two representative counterfactuals, routing lowers Qwen3-32B cost by 37.8% and raises DeepSeek-V3.2 average throughput by about 90% relative to direct official access. These results suggest that open-weight LLM deployment should be studied as a constrained statistical decision problem over a heterogeneous service layer, rather than as a static catalog of model capabilities.

Problem

Research questions and friction points this paper is trying to address.

open-weight LLMs

hosted APIs

service heterogeneity

model deployment

task-conditioned routing

Innovation

Methods, ideas, or system contributions that make the work stand out.

hosted LLM APIs

service heterogeneity

task-conditioned routing