Recent advances in large language models have made CLI-based AI agents a practical tool for accelerating GPU porting of large legacy scientific applications. Such applications, however, are not merely old code bases; they are scientific assets whose credibility has been accumulated through long-term development, comparison with observations, and use in domain studies. GPU porting must therefore pr
Validation-Centric AI-Assisted GPU Porting is a workflow designed to port large legacy scientific applications to GPUs while preserving their accumulated scientific validity. Using the CReSS weather simulation code (250,000+ lines) as a case study, researchers employed an AI agent to automate kernel extraction, OpenACC transformation, and rigorous dump-based validation.
The workflow achieved a 5.1x application-level speedup over the CPU baseline and produced numerically validated GPU implementations for 162 target kernels. Crucially, the validation-centric approach detected five numerical discrepancies caused by floating-point differences, enabling developers to address issues that might otherwise be hidden in full application runs. This demonstrates that AI-assisted porting requires not just code generation, but session-spanning context management and runtime-state reconstruction to ensure scientific credibility.
This material presents a case study in using CLI-based large-language-model agents to help port a very large legacy weather simulation codebase—over 250,000 lines—to GPU execution. The central framing is that such codes are not just legacy software to be modernized; they are scientific instruments whose value depends on long-accumulated validation, observational comparisons, and domain trust. A GPU port therefore cannot be judged only by compile success or speedup. It must preserve the numerical behavior, physical assumptions, and diagnostic outputs that make the model scientifically credible.
The key contribution is a validation-centric workflow for AI-assisted GPU porting, in which the LLM agent is used as a practical software-engineering assistant inside the terminal while correctness checks remain the primary gate. Rather than treating AI-generated or AI-suggested changes as automatically reliable, the approach embeds regression testing, output comparison, and domain-level validation into the porting loop. The paper’s broader insight is that AI can substantially reduce the mechanical burden of large-scale refactoring and GPU adaptation, but only if the process is organized around continuous verification of scientific fidelity.
This matters because it addresses a real bottleneck in HPC modernization: many critical scientific codes are too large, too domain-specific, and too deeply validated to be safely rewritten by hand in a short timeframe. By showing that AI agents can be used productively within a validation-first discipline, the work offers a plausible path for domain scientists to accelerate GPU migration without sacrificing the trustworthiness of long-running simulation assets. It also highlights an important research direction in AI-assisted scientific software engineering: moving beyond “can the model generate code?” to “can the generated changes be verified as scientifically equivalent?”