Training high-resolution AI-based Earth forecasting models is memory-intensive. Window-based Swin Transformers reduce the quadratic cost of global attention, but existing distributed systems such as AERIS primarily target pixel-level models and do not jointly support convolutional sampling modules and shifted-window execution. Long-lead rollout finetuning further increases activation memory. To ad
TERRA is a hierarchical parallel training and memory orchestration framework designed to address the memory-intensive nature of high-resolution AI-based Earth forecasting models. While window-based Swin Transformers reduce the quadratic cost of global attention, existing systems like AERIS primarily target pixel-level models, leaving a gap for hierarchical approaches that TERRA fills.
Key features of the TERRA framework include:
TERRA is a systems framework for training high-resolution, AI-based Earth forecasting models, particularly those built around window-based Swin-style transformers. The material highlights a systems gap: shifted-window attention reduces the quadratic cost of global attention over large spatial grids, but it introduces window-local, non-uniform execution patterns that are awkward for existing distributed training stacks. Convolutional sampling modules add further complexity, and long-lead rollout finetuning amplifies activation-memory pressure. Existing frameworks such as AERIS are described as being optimized mainly for pixel-level models, leaving insufficient support for jointly handling shifted-window execution, convolutional sampling, and the large activation footprints required by modern Earth-model training.
The core contribution is a hierarchical parallel training and memory-orchestration approach that co-designs distributed execution with activation-memory management. Instead of treating parallelism and memory control as separate optimizations, TERRA coordinates them across multiple levels of the training workload so that high-resolution spatial data, window-based transformer operations, and long rollout sequences can be executed within accelerator memory limits. The result is an architecture-aware runtime that makes windowed Earth-model training more scalable and practical on constrained GPU clusters.
This matters because high-resolution, long-horizon Earth forecasting is increasingly important for AI-based weather and climate prediction, but its practical deployment is bottlenecked by memory, not only compute. By addressing the joint execution and orchestration challenges of shifted-window transformers and related Earth-modeling components, TERRA lowers the barrier to training larger, higher-resolution, and longer-lead forecasting models, with potential benefits for both research-scale model development and operational AI Earth systems.