arXiv:2609.29395v1 Announce Type: new Abstract: Training with a fixed global batch limits how many distributed clients can provide examples in any one step. We examine a way to use additional server workers without increasing the batch processed by an individual workload. Global Clustered Parallel Split Learning (GCPSL) assigns clients to fixed clusters, executes a Parallel Split Learning with Gl

Topological visualization of Concurrent Split Learning Through Stable Client Clustering
Brave API

Concurrent Split Learning Through Stable Client Clustering (arXiv:2609.29395v1), submitted on September 24, 2026, introduces Global Clustered Parallel Split Learning (GCPSL) to overcome the participation limits imposed by fixed global batch sizes in split learning.

  • Mechanism: The method assigns clients to fixed, stable clusters, executing Parallel Split Learning with Global Sampling (GPSL) workloads for each cluster concurrently on separate server workers.
  • Synchronization: Client and server model segments are periodically fused at epoch boundaries, allowing concurrent training without moving raw data.
  • Performance: On a four-H100 GPU setup, label-aware GCPSL achieves 85% CIFAR-10 validation accuracy in 6.13 minutes, which is 3.11× faster than serializing the same workloads.
  • Trade-offs: While increasing the number of clusters improves direct data participation (e.g., size-balanced clustering increases participation by 3.25 percentage points), smaller clusters may incur an accuracy cost due to reduced label diversity.
Generated 7d ago
Open-Weights Reasoning

This paper addresses a scaling bottleneck in distributed split learning: when the global batch size is fixed, only a limited number of client examples can be incorporated into each training step, even if additional server-side compute is available. The authors propose Global Clustered Parallel Split Learning (GCPSL), a system-level approach that uses multiple server workers concurrently without increasing the batch size processed by any individual workload. In other words, the goal is to improve training throughput and server utilization while preserving the per-worker computational and memory constraints that are often critical in practical deployments.

The central mechanism is stable client clustering. Rather than dynamically reassigning clients in a way that can create coordination overhead or load imbalance, GCPSL assigns clients to fixed clusters that are mapped to server workers. Each cluster then participates in a parallel split-learning process, with global coordination across the workers. The stability of the clustering is important: it reduces the cost of reconfiguration, supports more predictable communication patterns, and allows the system to scale the number of participating clients across workers without changing the effective batch size seen by each worker.

The work matters because it reframes the scaling problem in split learning as a scheduling and clustering problem rather than simply a model-architecture problem. For privacy-sensitive or edge-aware settings where raw data must remain local and training is split between client and server components, the ability to add server workers without inflating per-worker batches can improve throughput, lower latency, and make distributed training more practical. More broadly, the paper contributes a useful systems perspective on how stable grouping of clients can enable concurrent, scalable split learning under fixed batch-size constraints.

Generated 7d ago
Sources