The Inference Bottleneck: A Formal Model of Vertical Foreclosure in AI Markets

📅 2026-04-19
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This study addresses the risk of non-price vertical foreclosure in the commercialization of generative AI, particularly within inference, distribution, and routing layers, which may undermine market competition. It presents the first formal model capturing mechanisms such as quality-of-service discrimination in AI inference markets, routing bias at the assistant layer, and tiered access restrictions. Integrating game-theoretic analysis, Logit demand functions, and symmetric competition assumptions, the model is calibrated using projected 2026 data from four major providers. The analysis identifies equilibrium conditions for service-quality gaps and their boundary relationships with key market parameters. The paper proposes a “neutral inference” regulatory framework, with quantitative results indicating that Google and OpenAI possess the strongest foreclosure capabilities, while Microsoft’s multi-model strategy constrains its leverage. Implementation of the proposed framework could generate annual net welfare gains amounting to tens of billions of dollars.

Technology Category

Application Category

📝 Abstract
As generative AI commercializes, competitive advantage is shifting from model training toward inference, distribution, and routing. This paper develops a formal game-theoretic model of vertical foreclosure in inference markets, as the formal-model companion to Besanson and Celani (2026). The model isolates two foreclosure mechanisms operating without predatory pricing: quality-of-service (QoS) discrimination against downstream rivals via latency, throughput, context limits, or feature access; and routing bias in assistant-layer interfaces. An extension motivated by Anthropic's April 2026 release of Claude Opus 4.7 alongside the restricted-access Claude Mythos Preview introduces a third mechanism, tier-based access discrimination, parameterized by a tier gap (tau) and partner-exclusivity (kappa). The main result gives an explicit local equilibrium characterization of the QoS gap. Under logit demand and symmetric rivals, the gap is strictly increasing in inference-quality importance (alpha) and downstream margins, and strictly decreasing in API price and rival entry elasticity. Discrimination vanishes at a joint boundary rather than at a simple threshold in alpha alone. A stylized calibration to four providers using April 2026 data treats parameter values as inputs to a comparative risk mapping, not structural estimates. The mapping suggests Google and OpenAI face conditions most conducive to foreclosure; Microsoft's realized routing bias has been voluntarily constrained by a March 2026 multi-model pivot; Anthropic shows low consumer-channel risk and elevated risk in enterprise coding-agent segments. The policy section proposes Neutral Inference, a four-pillar conduct framework: QoS parity, routing transparency, FRAND-style non-discrimination, and tier transparency with release-pathway discipline. Illustrative welfare calculations suggest net gains in the tens of billions annually.
Problem

Research questions and friction points this paper is trying to address.

vertical foreclosure
inference markets
quality-of-service discrimination
routing bias
tier-based access
Innovation

Methods, ideas, or system contributions that make the work stand out.

vertical foreclosure
quality-of-service discrimination
routing bias
tier-based access
Neutral Inference
🔎 Similar Papers