Growing demand for artificial intelligence (AI) inference services requires scalable infrastructure, yet centralized serving costs rise with demand. We propose a collaborative distributed inference system combining dedicated infrastructure with resources contributed by service users. Dedicated resources provide baseline capacity for maintaining quality of service (QoS), while volunteered resources
The search context does not contain information about a specific system named "User-Assisted Collaborative Distributed Inference." However, research indicates that collaborative distributed inference architectures optimize QoS and costs by distributing model execution across device edge, far edge (5G MEC), and cloud tiers.
Key strategies for scalable, QoS-aware autoscaling include:
This paper addresses the cost and scalability challenges of serving AI inference workloads by proposing a user-assisted collaborative distributed inference architecture. Rather than relying solely on centralized, dedicated infrastructure, the system combines a baseline pool of managed resources with compute contributed by service users themselves. The dedicated infrastructure is intended to preserve a minimum level of quality of service (QoS), while volunteered resources provide additional elasticity for overflow or peak demand. In effect, the work reframes AI inference autoscaling as a hybrid resource-orchestration problem: some capacity is provisioned for guaranteed performance, while other capacity is opportunistically recruited from a heterogeneous user base.
A central contribution is a QoS-aware autoscaling approach that treats user-contributed resources as a complementary elasticity layer rather than a drop-in replacement for dedicated servers. The system must decide how to route requests, when to admit or shed user resources, and how to maintain latency, throughput, or reliability targets in the face of variability in volunteered compute. The key insight is that inference demand is often bursty and not uniformly critical, so a portion of the workload can be served on less controlled or less reliable resources without violating overall service objectives.
The material matters because it offers a potential path toward more cost-efficient AI serving infrastructure. As inference demand grows, purely centralized scaling can become expensive and capacity-constrained, especially during traffic spikes. By leveraging user-side resources, the proposed approach can reduce peak infrastructure costs, improve utilization, and support more sustainable autoscaling for AI services. It is particularly relevant to settings where edge devices, consumer hardware, or distributed compute pools can participate in serving models while still respecting QoS requirements.