Introduces probabilistic alignment as an evaluation criterion requiring video world models to reproduce the full distribution of valid behaviors rather than single plausible trajectories.

Topological visualization of PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
Brave API

PAWBench is a diagnostic benchmark that evaluates video generators as stochastic samplers of world dynamics, requiring them to reproduce the full distribution of valid behaviors rather than just individual plausible trajectories. It defines probabilistic alignment through two criteria: probability-mass alignment (matching reference frequencies of outcomes) and support alignment (recovering the range of physically possible futures).

The benchmark tests 11 current video generation models across 50 scenarios using PAWEval, an outcome-level protocol that converts repeated rollouts into empirical distributions. Results indicate that no model consistently matches reference probabilities while recovering valid futures, with even the best-performing model (Cosmos 3 Super I2V) failing to achieve accurate calibration across all scenes.

Interventions via language prompting, noise sampling, and fine-tuning were found insufficient to reliably reshape the model’s intrinsic predictive distribution to match scene-conditioned probabilities. This highlights that plausible, diverse, or controllable single-rollouts do not establish true probabilistic alignment, as current models often concentrate probability mass on narrow subsets of futures rather than reflecting the underlying stochastic structure.

Generated Aug 30, 2026
Open-Weights Reasoning

PAWBench reframes the evaluation of video world models around probabilistic alignment: a model should not merely generate a single visually plausible continuation of an observed sequence, but should reproduce the full conditional distribution over valid future behaviors. The paper argues that current video-generation and world-modeling benchmarks often reward perceptual realism or single-trajectory accuracy, which can mask serious failures in distributional fidelity. A model may produce convincing frames while assigning probability mass to implausible futures, neglecting valid alternatives, or being poorly calibrated about uncertainty.

The key contribution is a benchmark and evaluation criterion for measuring how well world models match the distribution of possible futures, rather than only their average visual quality. By focusing on coverage, calibration, and alignment with a target distribution of valid behaviors, PAWBench can separate models that are sharp but underdiverse from those that are diverse but poorly calibrated. This makes it possible to diagnose whether a model is learning the structure of the environment or merely learning to produce plausible-looking samples.

The work matters because many downstream uses of world models—robotics, planning, policy evaluation, counterfactual reasoning, and interactive simulation—depend on more than plausible video generation. They require reliable estimates of which futures are likely, which are merely possible, and where the model is uncertain. In that sense, PAWBench shifts the question from “Can the model generate a good video?” to “Does the model understand the probabilistic structure of the world it is modeling?”

Generated Aug 30, 2026
Sources