Model Cards and Data Cards have demonstrated the value of structured, human-readable documentation for machine learning artifacts, capturing their context, parameters, limitations, and intended use. However, these practices remain focused on static artifacts (the datasets and trained models themselves) while overlooking the workflow executions that produce, transform, and evaluate them. Such execu
Workflow Cards are a new documentation framework that provides structured, human- and LLM-readable summaries of workflow execution provenance, addressing the gap left by static Model Cards and Data Cards.
This material proposes Workflow Cards, a documentation practice for capturing the execution history of machine learning workflows using provenance data. It builds on the success of Model Cards and Data Cards, which provide structured, human-readable descriptions of static ML artifacts such as trained models and datasets. The paper identifies a gap in this approach: much of the behavior, risk, and context of an ML system is not contained in the final artifact alone, but in the sequence of operations that produced it—data ingestion, preprocessing, training, evaluation, deployment, and downstream use. Workflow Cards extend structured documentation to these dynamic execution processes, treating the workflow run itself as a first-class object worthy of explanation.
The central contribution is a framework for turning provenance records into concise, interpretable summaries of workflow executions. Rather than only describing what a model or dataset is, a Workflow Card records what happened during an execution: the inputs and versions used, the transformations applied, the configuration and parameters involved, the intermediate and final outputs, evaluation results, and relevant contextual or limitation information. This makes the card a bridge between low-level provenance data and the kind of high-level explanation needed for reproducibility, debugging, governance, and stakeholder communication.
The work matters because modern ML systems are increasingly pipeline-based, and failures or undesirable behaviors often depend on execution context rather than on the final model alone. By documenting workflow executions in a structured way, Workflow Cards can help teams trace how data or code changes affect outcomes, compare runs, audit model behavior, and identify sources of bias, instability, or non-reproducibility. More broadly, the paper positions provenance not just as a technical logging mechanism, but as a foundation for explainable, accountable, and reusable ML engineering practices.