arXiv:2609.20359v1 Announce Type: new Abstract: The symbiotic scaling of artificial intelligence models and high-performance computing systems continually creates algorithmic challenges in their convergence. Foundation models (FMs) are a crucial example, requiring months-long training on thousands of cutting-edge GPUs. Sharded data parallelism (DP) is the dominant strategy to accelerate such comp
Mittone and Aldinucci present FL+FSDP and FL+HSDP, hybrid algorithms that integrate Federated Learning (FedAvg) with Sharded Data Parallelism to reduce communication overhead in large-scale Foundation Model training. By decoupling GPU clusters into loosely-coupled federation groups, these methods limit global batch size growth and minimize traffic on slow interconnects.
Evaluated on a Llama3.1 8B model across 512 A100 GPUs, the approaches achieve up to 8.04× faster data processing and 4.48× lower evaluation perplexity compared to standard FSDP and HSDP. This improvement stems from reduced communication costs and more stable convergence trajectories enabled by periodic, lightweight aggregations. The work was awarded the Best Paper Award at Euro-Par 2026.
Accelerating Sharded Data Parallelism at Scale with Federated Learning
This paper addresses a core scaling problem in foundation-model training: as AI models and high-performance computing systems grow together, sharded data parallelism (DP)—the dominant strategy for training across large GPU clusters—faces increasingly difficult convergence behavior. The abstract frames the issue as an algorithmic, not merely hardware-level, challenge: training foundation models can require months of compute on thousands of accelerators, and simply scaling synchronous DP may not preserve efficient convergence as system size, model size, and coordination overhead grow.
The central contribution is to apply federated learning (FL) concepts to sharded DP at datacenter scale. Rather than treating FL only as a privacy-preserving or edge-training paradigm, the work appears to use FL-style coordination mechanisms to improve how distributed updates are computed, combined, or synchronized across many workers. The key insight is that techniques developed for federated settings—such as local update computation, flexible participation, and relaxed or asynchronous aggregation—may help reduce coordination bottlenecks and improve the convergence behavior of large-scale DP training.
This matters because foundation-model training is increasingly limited by the interplay between optimization, communication, and cluster-level performance. If FL-inspired methods can improve the throughput-to-convergence tradeoff in sharded DP, they could reduce wall-clock training time, increase robustness to load imbalance or stragglers, and make large-scale training more efficient as GPU clusters continue to expand. More broadly, the work bridges distributed systems, optimization, and ML infrastructure, suggesting that federated techniques may have first-order value for high-performance AI training, not just privacy-sensitive or decentralized settings.