User-Assisted Collaborative Distributed Inference for Efficient QoS-Aware Autoscaling

📅 2026-08-12
📈 Citations: 0
Influential: 0
📄 PDF
🤖 AI Summary
This work addresses the escalating costs of centralized AI inference deployments, which struggle to simultaneously meet quality-of-service (QoS) requirements and support elastic scaling. To overcome this challenge, the paper proposes the first collaborative distributed inference architecture that synergistically integrates user-contributed volunteer resources with dedicated infrastructure. It introduces a structured temporal factorization-based high-dimensional generative Markov model to capture dynamic stochastic interactions, enabling joint QoS-aware optimization of task scheduling and resource allocation. Extensive simulations demonstrate that, as user scale grows, the proposed approach significantly improves request completion rates, reduces P99 latency, and substantially decreases reliance on dedicated resources—thereby validating its feasibility in enhancing infrastructure efficiency while guaranteeing QoS.
📝 Abstract
Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources absorb increasing demand without proportional growth in centralized infrastructure. To capture stochastic and dynamic interactions among users, resources, tasks, and policies, we develop a high-dimensional generative Markov model with structured temporal factorization. The model supports simulation and provides a foundation for task scheduling and QoS-aware resource allocation optimization. We evaluate the system across user populations, resource capacities, and centralized and distributed scheduling policies. Simulations show that distributed scheduling becomes increasingly advantageous as the user population grows, improving request completion and P99 latency while substantially reducing dedicated resource consumption. These results demonstrate the feasibility of user-assisted collaborative inference for infrastructure-efficient autoscaling.
Problem

Research questions and friction points this paper is trying to address.

distributed inference
QoS-aware autoscaling
collaborative computing
infrastructure efficiency
AI inference services
Innovation

Methods, ideas, or system contributions that make the work stand out.

collaborative distributed inference
QoS-aware autoscaling
generative Markov model
user-assisted computing
structured temporal factorization
🔎 Similar Papers
No similar papers found.