arXiv:2609.20497v1 Announce Type: new Abstract: We introduce Adaptive Score-based Routing Balancer (ASRB), a dynamic, score-based request routing mechanism for Kubernetes-based service deployments over the computing continuum. ASRB jointly considers infrastructure-level information, response time measurements, and application-level quality indicators, with a particular focus on serving Machine Le

Topological visualization of A Kubernetes-Native Request Router for Quality-Aware Inference Serving in the Computing Continuum
Brave API

ASRB (Adaptive Score-based Routing Balancer) is a Kubernetes-native request router designed to optimize quality-aware inference serving in the computing continuum. It dynamically balances requests across service instances by jointly considering infrastructure metrics, response times, and application-level quality indicators, allowing operators to flexibly manage latency-accuracy trade-offs via weighted policies.

Key performance characteristics include: Reduced Latency: Achieves at least 10 ms lower mean response time compared to state-of-the-art latency-oriented routers. Lower Overhead: Reduces monitoring costs by up to ~70% through selective, adaptive monitoring intensity without sacrificing routing quality. * Operational Efficiency: Requires no modifications to Kubernetes, enabling straightforward deployment in existing clusters while maintaining reduced failure rates and high responsiveness to environmental changes.

Generated 15d ago
Open-Weights Reasoning

This paper presents Adaptive Score-based Routing Balancer (ASRB), a Kubernetes-native request routing mechanism designed for ML inference serving across heterogeneous computing continuum environments. Rather than relying solely on static load-balancing policies or simple resource-utilization thresholds, ASRB combines infrastructure-level signals, observed response-time behavior, and application-level quality indicators into a dynamic scoring model. The router uses these scores to direct incoming inference requests toward service replicas or nodes that are expected to satisfy both performance and quality objectives, making it particularly relevant for latency-sensitive or quality-dependent ML workloads.

The key insight is that effective inference serving in the computing continuum requires routing decisions that reflect runtime quality, not just capacity. Inference services can appear healthy from a resource standpoint while still degrading in response time, tail latency, or application-specific quality due to workload skew, model complexity, network conditions, or node heterogeneity. By embedding quality-aware scoring into the routing layer, ASRB aims to improve request placement under dynamic conditions while remaining compatible with existing Kubernetes deployment patterns.

The work matters because it addresses a practical gap between conventional cloud/edge load balancing and the needs of ML inference systems. As inference workloads are increasingly distributed across edge, on-prem, and cloud resources, operators need mechanisms that can adapt to changing service quality without requiring application-specific redesign. A Kubernetes-native, score-based router like ASRB provides a potentially reusable control-plane component for improving SLO adherence, resource efficiency, and service reliability in continuum-based inference deployments.

Generated 15d ago
Sources