arXiv:2610.00971v1 Announce Type: new Abstract: Moving a workflow developed and tested at one HPC site to another rarely succeeds without some amount of trial and error. Package managers rebuild software environments, containers ship whole filesystems, and workflow specifications such as backpacks package a workflow with its software, data, and resource requirements. These approaches address one
The Other Half of Workflow Portability proposes the HPC site profile, a structured, evidence-backed document that captures how a specific High-Performance Computing (HPC) site must be used, addressing the gap left by tools that only package workflow requirements. Authored by Md Saiful Islam and Douglas Thain from the University of Notre Dame, this work was accepted to WORKS 2026 and submitted to arXiv on October 1, 2026.
The authors automate profile construction through a three-step pipeline: probing the login node for measurable configuration, using a bounded language model agent to extract typed fields from authoritative documentation, and submitting pilot jobs to verify compute-node behaviors and network policies. Every field is subject to deterministic verification, ensuring that only values supported by evidence enter the profile, which then enables a preflight step to generate execution plans or identify incompatibilities before workflow submission.
Evaluation across sites including Purdue Anvil, TACC Stampede3, and the Notre Dame CRC demonstrated high precision (0.957–0.992) and citation integrity (1.000), with standard BM25 retrieval achieving comparable quality to more expensive strategies. The open-source implementation, schemas, and detailed results are available at the floability/hpc-site-preflight GitHub repository.
The paper frames HPC workflow portability as a two-sided problem. Existing mechanisms—package managers, containers, and workflow packaging formats such as backpacks—primarily make the workflow side portable by bundling code, dependencies, data, and resource requirements. The authors argue that this is only half of the solution: a workflow also has to be compatible with the destination site’s software availability, filesystems, schedulers, storage, network topology, authentication model, and other operational constraints. Without an explicit, verifiable account of the target site, users still face ad hoc trial and error when moving workloads across HPC environments.
Its central contribution is an “evidence-backed” HPC site profile, paired with an agentic discovery process for populating and validating that profile. Rather than treating site metadata as static or manually curated, the approach uses agents to discover relevant site characteristics and to ground profile claims in evidence—such as documentation, site metadata, or live probes—so that the resulting profile is more reliable than a simple self-description. The key insight is that site compatibility should be represented as a machine-readable, inspectable artifact that can be compared against a workflow’s requirements, making portability failures easier to diagnose before execution.
This matters because it shifts HPC portability from “ship the environment” toward “understand the destination.” Such profiles could support better site selection, pre-flight validation, multi-site execution planning, and reproducible workflow deployment, reducing the manual debugging that currently dominates HPC workflow migration. In a broader sense, the work complements container- and package-based portability by addressing the site-side half of the problem, which is especially important as workflows increasingly span heterogeneous clusters, cloud systems, and shared research infrastructure.