arXiv:2609.20723v1 Announce Type: new Abstract: Online image generation with Diffusion Transformers (DiTs) must meet latency service-level objectives (SLOs) while using GPU resources efficiently. Existing systems improve GPU utilization by batching multiple requests for joint execution. However, request-level batching offers limited control over batch size: batches may be too small to saturate GP
PixelFlow is a distributed DiT serving system that improves efficiency by managing workloads at the token level rather than the request level, allowing GPUs to aggregate fragmented residual capacity to serve mixed-resolution requests. By using image tokens as a common unit for scheduling, it enables fine-grained resource sharing that reduces queueing delays and increases concurrency without violating latency Service Level Objectives (SLOs).
Key performance metrics from evaluations on H100 GPUs with Stable Diffusion 3 and FLUX.1-dev include: Up to 43% improvement in SLO attainment compared to state-of-the-art systems. Up to 2.8× higher goodput, significantly outperforming existing solutions like TetriServe. * Enhanced GPU utilization by dynamically adapting to heterogeneous workloads through synchronization groups that isolate SLO-critical requests while sharing idle capacity.
PixelFlow addresses a systems bottleneck in serving diffusion-transformer (DiT) image generators: how to meet latency service-level objectives (SLOs) while keeping GPUs sufficiently utilized. The paper argues that conventional request-level batching is too coarse for DiT workloads. A DiT request is not a single monolithic compute unit; it consists of many denoising steps, each operating over large token sequences. As a result, batching whole requests gives limited control over the effective batch size—small batches underutilize the GPU, while large batches can risk SLO violations or create uneven load across distributed devices.
The key contribution is a token-level workload-management approach for distributed DiT serving. Rather than scheduling entire image requests as atomic jobs, PixelFlow decomposes inference work into finer-grained token-level units and schedules those units across GPUs and time. This allows the system to pack work more tightly, rebalance load as request arrivals or denoising-stage costs vary, and make scheduling decisions based on latency constraints rather than only on request boundaries. In effect, PixelFlow treats DiT serving as a fine-grained workload-management problem, not merely a continuous-batching or model-parallelism problem.
The work matters because it targets a central cost/latency tradeoff in production image generation. Large DiTs are compute-intensive, and providers need high throughput without sacrificing user-facing latency. By moving from request-level to token-level scheduling, PixelFlow can improve GPU utilization and SLO adherence on the same hardware, making distributed DiT serving more efficient and economically viable. More broadly, the paper highlights a useful principle for transformer-based generative systems: when inference work is highly regular and parallelizable, finer-grained workload management can unlock efficiency that coarse request-level serving leaves on the table.