Large language model (LLM) agents are increasingly used to modernize the legacy Fortran underlying production scientific software, but validation of these transformations emphasizes nominal executions and may not test whether a modernization preserves the original code's response to faults, perturbations, and reduced precision. We present a differential fault-injection validation method: a harness
Validating LLM-modernized scientific software through differential fault injection requires moving beyond nominal execution checks to ensure numerical fidelity and robustness. Recent research highlights two critical gaps: standard validation often ignores semantic-preserving mutations that test fault localization robustness, and legacy translations (e.g., Fortran to JAX) must be verified against Fortran reference oracles for both functional parity and gradient accuracy.
1. Numerical and Differentiability Parity For scientific code modernization (e.g., translating Fortran to JAX), validation must enforce module-level numerical parity using Fortran reference inputs/outputs before integration. The translated code must not only reproduce physics to high precision but also support gradient-based parameter estimation, requiring comprehensive reverse-mode gradient verification through coupled physics columns to ensure autodiff accuracy.
2. Differential Fault Injection and Robustness Standard LLM evaluation often lacks rigorous fault injection frameworks. MAS-FIRE introduces a systematic framework for Multi-Agent Systems that injects 15 fault types (e.g., prompt modification, response rewriting) to evaluate reliability. For code-specific validation, researchers propose Semantic-Preserving Mutations (SPMs) to test how LLMs handle changes that preserve program behavior but alter structure, revealing weaknesses in fault localization that nominal tests miss.
3. Behavioral Equivalence via Feedback Loops Modernization frameworks like AgentModernize emphasize that iterative feedback loops are essential for preserving business logic. Without an Equivalence Validator agent that detects behavioral divergences and feeds corrections back to the transformer, generated code may compile but fail gold-standard behavioral tests. This approach uses Behavioral Specification Graphs to capture constraints and performs differential trace analysis to verify equivalence.
To ensure reliable modernization, validation must integrate static dependency analysis, multi-agent fault injection, and statistical effect size validation to confirm that modernized scientific software maintains both numerical accuracy and operational robustness.
This material addresses a validation gap in LLM-driven modernization of legacy scientific software, particularly Fortran codebases that underpin production scientific workflows. LLM agents can perform large-scale transformations—such as language modernization, refactoring, or porting—but the paper argues that conventional validation, which often focuses on nominal executions and matching expected outputs, is insufficient. Such testing may confirm that a modernized program produces the same results on clean inputs while missing whether it preserves the original software’s behavior under faults, environmental perturbations, or reduced numerical precision.
The proposed approach is a differential fault-injection validation method, implemented as a harness that compares the behavior of an original legacy code and its LLM-modernized counterpart under controlled stress conditions. Rather than treating correctness as equivalence only on normal runs, the method injects faults or perturbations into both versions and evaluates whether their responses remain aligned. This shifts the validation question from “does the modernized code produce the same nominal output?” to “does it preserve the original code’s failure modes, numerical sensitivity, and robustness characteristics?”
The contribution matters because scientific software often operates in numerically sensitive, long-running, and safety- or mission-critical contexts, where small semantic changes introduced by automated refactoring can have outsized consequences. By providing a systematic way to test behavioral equivalence beyond golden-output comparisons, the work offers a practical mechanism for increasing confidence in LLM-generated modernizations. More broadly, it highlights that trustworthy AI-assisted code transformation requires validation that accounts for edge-case behavior, not just functional correctness under ideal conditions.