arXiv:2609.19972v1 Announce Type: cross Abstract: Federated Learning (FL) is experiencing a substantial research interest, with many frameworks being developed to allow practitioners to build federations easily and quickly. Most of these efforts do not consider two main aspects that are key to Machine Learning (ML) software: customizability and performance. This research addresses these issues by
Efficiently Distributed Federated Learning (arXiv:2609.19972v1), published on September 17, 2026, by Gianluca Mittone, Robert Birke, and Marco Aldinucci from the University of Turin, introduces FastFederatedLearning (FFL), an open-source, high-performance C/C++ framework designed to address the customizability and performance limitations of existing Python-based FL systems.
FFL leverages the FastFlow parallel programming framework and PyTorch C++ interface to enable custom communication graphs and support for traditional ML models beyond Deep Neural Networks. In tests across heterogeneous hardware (x86-64, ARM-v8, RISC-V), FFL achieved 2.5x to 3.69x speedups compared to Intel OpenFL, while aiming to support dynamic federations and asynchronous communications for improved scalability in real-world distributed environments.
This material addresses a practical gap in Federated Learning (FL) systems: many existing frameworks make it easy to assemble federations, but they often underemphasize two properties that are central to real ML software—customizability and performance. The paper argues that building an FL federation quickly is insufficient if the resulting system cannot be adapted to diverse training objectives, model architectures, data pipelines, or deployment constraints, or if it introduces excessive communication, coordination, or computation overhead.
The central contribution is an approach to efficiently distributed federated learning that treats FL not merely as a privacy-preserving training paradigm, but as a distributed ML system design problem. By focusing on how federated workloads are distributed across participating nodes, the work aims to support flexible, user-defined training flows while reducing the runtime costs associated with federation. In effect, it seeks to preserve the adaptability expected of modern ML stacks without paying a prohibitive performance penalty for the added distribution layer.
This matters because FL is increasingly being moved from prototype experiments toward production-scale ML workloads, where heterogeneity, resource constraints, and communication costs are first-order concerns. A framework or methodology that improves both customizability and efficiency can make federated training more practical for large, distributed, and privacy-sensitive environments. The work is therefore relevant to practitioners and researchers who need FL systems that are not only easy to build, but also performant enough to support real-world ML operations.