arXiv:2609.09160v2 Announce Type: replace Abstract: We present LBFAST, a GPU-oriented lattice Boltzmann solver based on a lightweight moment-represented formulation, in which post-collision populations are reconstructed on the fly from a reduced set of moments rather than stored explicitly. This approach significantly lowers the memory footprint, enabling large three-dimensional simulations withi

Topological visualization of LBFAST: A Lightweight Moment-Represented Lattice Boltzmann Solver for Multi-GPU Architectures
Brave API

LBFAST is a GPU-oriented lattice Boltzmann solver that significantly reduces memory footprint by reconstructing post-collision populations on the fly from a reduced set of hydrodynamic moments rather than storing them explicitly. This lightweight, moment-represented formulation allows for large three-dimensional simulations within the constraints of modern accelerator architectures, achieving near-ideal weak scaling up to 512 GPUs and a ~47-63% reduction in memory usage compared to standard population-based implementations. The solver utilizes a velocity–pressure framework coupled with a conservative Allen–Cahn equation for phase fields, implemented in CUDA Fortran with MPI for distributed-memory parallel execution.

Generated 2d ago
Open-Weights Reasoning

LBFAST is a GPU-oriented lattice Boltzmann method (LBM) solver that replaces explicit storage of post-collision distribution functions with a lightweight moment-represented formulation. In conventional LBM implementations, especially for three-dimensional stencils, the solver must store and stream a large number of population values per lattice node, which makes memory capacity and memory bandwidth major bottlenecks on modern accelerators. LBFAST instead retains a reduced set of moments and reconstructs the post-collision populations on the fly during the update step, thereby reducing the persistent state that must reside in GPU memory and be communicated across devices.

The key contribution is the combination of this moment-based representation with a multi-GPU architecture tailored to the resulting data access patterns. The central insight is that much of the LBM state can be compressed into a smaller set of physically meaningful quantities, shifting part of the computational cost from large-scale memory traffic to comparatively lightweight reconstruction and streaming operations. This design is aimed at large-scale three-dimensional simulations, where explicit population storage would otherwise limit domain size, resolution, or achievable parallel efficiency.

The work matters because it addresses one of the practical limits of LBM in high-performance computing: although the method is attractive for complex, highly parallel fluid simulations, its memory footprint can become prohibitive on GPUs, particularly for high-resolution 3D problems. By lowering memory demand and improving suitability for multi-GPU execution, LBFAST makes it more feasible to run larger, longer, or higher-resolution lattice Boltzmann simulations while retaining the method’s natural parallelism and local update structure. This is especially relevant for applications in computational fluid dynamics where memory-bound solvers often constrain scalability more than raw floating-point throughput.

Generated 2d ago
Sources