Large language model (LLM) agents are increasingly used to modernize the legacy Fortran underlying production scientific software, but validation of these transformations emphasizes nominal executions and may not test whether a modernization preserves the original code's response to faults, perturbations, and reduced precision. We present a differential fault-injection validation method: a harness

Topological visualization of Validating LLM-Modernized Scientific Software Through Differential Fault Injection
Brave API

Validating LLM-modernized scientific software through differential fault injection requires moving beyond nominal execution checks to ensure numerical fidelity and robustness. Recent research highlights two critical gaps: standard validation often ignores semantic-preserving mutations that test fault localization robustness, and legacy translations (e.g., Fortran to JAX) must be verified against Fortran reference oracles for both functional parity and gradient accuracy.

Key Validation Challenges and Approaches

1. Numerical and Differentiability Parity For scientific code modernization (e.g., translating Fortran to JAX), validation must enforce module-level numerical parity using Fortran reference inputs/outputs before integration. The translated code must not only reproduce physics to high precision but also support gradient-based parameter estimation, requiring comprehensive reverse-mode gradient verification through coupled physics columns to ensure autodiff accuracy.

2. Differential Fault Injection and Robustness Standard LLM evaluation often lacks rigorous fault injection frameworks. MAS-FIRE introduces a systematic framework for Multi-Agent Systems that injects 15 fault types (e.g., prompt modification, response rewriting) to evaluate reliability. For code-specific validation, researchers propose Semantic-Preserving Mutations (SPMs) to test how LLMs handle changes that preserve program behavior but alter structure, revealing weaknesses in fault localization that nominal tests miss.

3. Behavioral Equivalence via Feedback Loops Modernization frameworks like AgentModernize emphasize that iterative feedback loops are essential for preserving business logic. Without an Equivalence Validator agent that detects behavioral divergences and feeds corrections back to the transformer, generated code may compile but fail gold-standard behavioral tests. This approach uses Behavioral Specification Graphs to capture constraints and performs differential trace analysis to verify equivalence.

Critical Gaps in Current Practices

  • Lack of Statistical Rigor: Many studies report p-values without effect size analysis (e.g., Cohen’s d), failing to distinguish statistically detectable changes from transformative practical significance.
  • Black-Box Evaluation: Existing frameworks often treat LLMs as opaque, lacking systematic transformation pattern taxonomies that explain why specific modernizations succeed or fail.
  • Absence of Functional Testing: High success rates in syntactic or metric-level improvements (e.g., 97.9%) do not guarantee verified behavioral equivalence or security, as many evaluations skip automated test suite execution.

To ensure reliable modernization, validation must integrate static dependency analysis, multi-agent fault injection, and statistical effect size validation to confirm that modernized scientific software maintains both numerical accuracy and operational robustness.

Generated Sep 1, 2026
Open-Weights Reasoning

This material addresses a validation gap in LLM-driven modernization of legacy scientific software, particularly Fortran codebases that underpin production scientific workflows. LLM agents can perform large-scale transformations—such as language modernization, refactoring, or porting—but the paper argues that conventional validation, which often focuses on nominal executions and matching expected outputs, is insufficient. Such testing may confirm that a modernized program produces the same results on clean inputs while missing whether it preserves the original software’s behavior under faults, environmental perturbations, or reduced numerical precision.

The proposed approach is a differential fault-injection validation method, implemented as a harness that compares the behavior of an original legacy code and its LLM-modernized counterpart under controlled stress conditions. Rather than treating correctness as equivalence only on normal runs, the method injects faults or perturbations into both versions and evaluates whether their responses remain aligned. This shifts the validation question from “does the modernized code produce the same nominal output?” to “does it preserve the original code’s failure modes, numerical sensitivity, and robustness characteristics?”

The contribution matters because scientific software often operates in numerically sensitive, long-running, and safety- or mission-critical contexts, where small semantic changes introduced by automated refactoring can have outsized consequences. By providing a systematic way to test behavioral equivalence beyond golden-output comparisons, the work offers a practical mechanism for increasing confidence in LLM-generated modernizations. More broadly, it highlights that trustworthy AI-assisted code transformation requires validation that accounts for edge-case behavior, not just functional correctness under ideal conditions.

Generated Sep 1, 2026
Sources