arXiv:2510.15596v3 Announce Type: replace Abstract: Large model training beyond tens of thousands of GPUs is an uncharted territory. At such scales, disruptions to the training process are not a matter of if, but a matter of when -- a stochastic process degrading training productivity. Dynamic runtime variation will become increasingly more frequent as training scales and GPUs are operated in inc
PRISM is a performance modeling framework for large-scale distributed training that treats runtime variability as a stochastic process rather than a deterministic one. Developed by researchers at Meta FAIR and Harvard University, it addresses the significant performance degradation caused by hardware and environmental variability in clusters exceeding 64,000 GPUs.
The framework utilizes a statistical method to quantify probabilistic guarantees on training time, specifically targeting the p95 execution time. By modeling operator-level latency distributions and propagating them through workload dependencies via Monte Carlo simulation, PRISM provides accurate performance estimates with a 5.4% error margin across diverse configurations.
Key findings and capabilities include: Observed Variability: At the 64k GPU scale, training exhibits 9-12% GPU time variability, with communication operations showing nearly 3× higher variability than compute operations. Optimization Insights: The model reveals that optimizing communication kernels like AllGather and ReduceScatter is most effective for minimizing training step time variability. Performance Gains: By accounting for sensitivity to variation, PRISM identifies up to 1.26× performance improvement potential through smarter computation node placement. Straggler Impact: A single node running at p95 latency can increase total training step time by 1.64× if placed poorly within the parallelization strategy.
PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training addresses a central operational challenge in frontier-scale AI: as training jobs span tens of thousands of GPUs, the dominant source of inefficiency is no longer algorithmic or hardware roofline limits alone, but stochastic runtime disruption. The paper frames failures, preemptions, stragglers, network perturbations, and other transient anomalies as a probabilistic process that continuously degrades training productivity. Its core contribution is a scalable modeling approach that moves beyond deterministic “best-case” performance estimates and instead characterizes expected runtime behavior under uncertainty, using runtime insights to build models that remain tractable at cluster scale.
The work is significant because it targets the regime where traditional performance engineering breaks down: large enough that rare events become frequent, and dynamic enough that static assumptions about availability, throughput, or checkpoint overhead are no longer reliable. By treating runtime variation as a first-class modeling object, PRISM enables more realistic predictions of effective training progress, productivity loss, and tail-risk behavior. This matters for capacity planning, fault-tolerance design, checkpointing policy, scheduling, and cost estimation, since decisions made under deterministic assumptions can be systematically optimistic in very large distributed training environments.
More broadly, the paper contributes to the emerging discipline of treating large-scale training systems as stochastic production systems rather than monolithic compute jobs. Its likely value is methodological: providing a principled way to quantify how runtime variability translates into end-to-end training performance, and to design systems that are resilient not just to individual faults, but to the aggregate statistical behavior of a massive, constantly perturbing training cluster. For technically literate readers, the key takeaway is that at the largest scales, performance modeling must be probabilistic, scalable, and grounded in runtime evidence if it is to support reliable operation of frontier training systems.