Long-context reasoning for large language models (LLMs) is becoming increasingly important, but training over long sequences remains challenging due to massive memory and communication requirements. Sequence parallelism has emerged as an essential technique for addressing bottlenecks in long sequence LLM training. However, we observe that existing sequence parallelism methods are batch-agnostic an
BASP (Batch-Aware Sequence Parallelism) is a training optimization for Large Language Models that reduces communication overhead by partitioning GPUs into disjoint groups based on micro-batch size. Unlike standard methods like Ulysses-SP that perform global all-to-all communication across all GPUs, BASP assigns each sequence in a batch to a smaller, independent subgroup of GPUs.
This approach decomposes a single large all-to-all collective into multiple smaller, parallel collectives, significantly reducing synchronization latency. Experimental results on NVIDIA A100 clusters demonstrate that BASP improves end-to-end training time by 1.17x to 1.31x for Llama and Qwen models while maintaining identical accuracy and memory usage.
Key Mechanisms and Benefits: Reduced Communication: Limits all-to-all participants from $N$ (total GPUs) to $K=N/B$ (where $B$ is batch size), lowering communication time by up to 3.10x. Scalability: Speedup increases with larger batch sizes and longer sequence lengths, achieving up to 25.9% faster step times at 32K context lengths. * Compatibility: Preserves per-GPU memory footprint and integrates with ZeRO-3 for additional memory savings.
Summary
The paper addresses a central scaling bottleneck in long-context LLM training: as sequence lengths grow, activation memory and inter-GPU communication become dominant costs, even when tensor and pipeline parallelism are already in use. Sequence parallelism is a natural remedy because it partitions activation tensors along the sequence dimension, but the authors argue that existing sequence-parallel schemes are largely batch-agnostic: they partition and synchronize each sequence with little regard for the composition, heterogeneity, or batching structure of the training workload. This can lead to unnecessary collective traffic, poor load balance under variable-length sequences, and underutilized communication bandwidth.
BASP (Batch-Aware Sequence Parallelism) reframes sequence parallelism as a batch-level scheduling and partitioning problem. Rather than treating each sequence in isolation, it co-designs sequence partitioning, device placement, and communication scheduling with the batch composition. The key insight is that sequence-parallel operations can often be organized, packed, or overlapped across batch elements, reducing redundant collective traffic and improving overlap between communication and computation. By making the parallelization strategy batch-aware, BASP targets both memory reduction and communication efficiency, particularly in regimes with long, irregular, or mixed-length sequences.
This matters because long-context and reasoning workloads are pushing LLM training into regimes where communication, rather than raw compute or parameter memory, is often the primary limiter. A batch-aware sequence-parallel method can improve effective throughput and scalability on multi-GPU or multi-node clusters while preserving the memory benefits of sequence partitioning. The contribution is therefore a systems-level algorithmic improvement that complements existing parallelism stacks and is especially relevant for training or fine-tuning models on long-context data.
BASP: Communication-Efficient Batch-Aware Sequence Parallelism for LLM Training
This paper addresses a key systems bottleneck in training large language models on long contexts: sequence parallelism is now widely used to distribute activation memory and attention computation across devices, but many existing implementations treat the batch dimension as merely a collection of independent samples. As a result, communication patterns are often optimized per sequence rather than across the full batch, leading to redundant or underutilized collectives, smaller message sizes, and poorer overlap between computation and communication. The paper frames this “batch-agnostic” behavior as a major source of inefficiency, especially in long-sequence regimes where communication volume and latency can dominate training throughput.
The central contribution is BASP, a batch-aware sequence parallelism strategy that coordinates sequence partitioning and communication across both the sequence and batch dimensions. Rather than issuing sequence-parallel operations independently for each batch item, BASP exploits batch-level structure to aggregate, schedule, and fuse communication more efficiently. The key insight is that the batch dimension is not just a training convenience; it is a resource for improving collective efficiency, reducing redundant traffic, and better matching communication granularity to available bandwidth. By making sequence parallelism batch-aware, the method aims to lower communication overhead while preserving the memory benefits needed for long-context training.
This matters because long-context capability is becoming a first-class requirement for LLMs, particularly for reasoning, retrieval-augmented generation, multi-turn interaction, and document- or video-scale modeling. Training at such lengths is often limited less by raw compute than by memory and inter-device communication. A communication-efficient sequence parallelism scheme like BASP can therefore improve hardware utilization, increase effective throughput, and make long-context training more practical on existing clusters. More broadly, the work highlights an underexplored axis of optimization in distributed LLM training: jointly considering sequence layout, batch organization, and collective communication rather than treating them as independent system choices.