arXiv:2609.33544v1 Announce Type: new Abstract: Federated learning (FL) is increasingly deployed as a managed learning service rather than as a set of isolated training jobs. In networked edge environments, dependent FL service flows must coordinate heterogeneous clients, non-IID data, fluctuating communication latency, and precedence-constrained tasks under service-level completion requirements.
A-CoDa is an adaptive clustered coordination framework designed to manage dependent Federated Learning (FL) workflows in heterogeneous edge networks, addressing challenges like non-IID data, fluctuating latency, and precedence-constrained tasks. It utilizes LDD-based greedy-balanced clustering to group clients by label-distribution divergence, ensuring statistical coherence and balanced cluster sizes.
The framework employs FedMIX, an uncertainty-aware mechanism that ranks clients using a loss–latency–uncertainty utility to adaptively control intra-cluster exploitation and inter-cluster probing. This is combined with a PPO-based DAG scheduler to coordinate layer-wise task execution, minimizing end-to-end completion time while maintaining competitive accuracy across tasks such as medical imaging and wearable sensing.
The paper examines federated learning (FL) as a managed, workflow-level service rather than a collection of independent training jobs. It focuses on edge-network settings where dependent FL service flows must be coordinated across heterogeneous clients with non-IID data, variable communication latency, and precedence-constrained tasks. The central problem is to satisfy service-level completion requirements while managing the operational complexity introduced by client heterogeneity, network fluctuation, and task dependencies.
Its key contribution is an adaptive client clustering and coordination approach for FL workflow management. By grouping clients according to relevant system and data characteristics, and by coordinating the execution of dependent tasks, the work aims to reduce communication overhead, mitigate the impact of stragglers, and improve the likelihood that each service flow completes within its deadline. The broader insight is that edge FL performance depends not only on model aggregation quality, but also on how effectively the system orchestrates clients, task ordering, and timing under real-world constraints.
This matters because production edge deployments of FL face practical challenges that standard FL benchmarks often underrepresent, including limited bandwidth, unstable links, heterogeneous devices, and strict service-level objectives. By framing FL as a managed workflow and emphasizing adaptive coordination, the material points toward a more operational model for edge-based learning services, where reliability, timeliness, and resource efficiency are as important as statistical model performance.