Identifies misalignment between supervised next-state prediction training of web-agent world models and downstream ranker needs for discriminative state predictions across candidate actions.
Discriminative World Models address the misalignment where standard supervised next-state prediction (SFT) optimizes for surface-level token accuracy rather than the discriminative ranking required to distinguish between candidate actions. While SFT-trained models may achieve high exact-match accuracy on predicted states, they often fail to preserve behavioral consistency, leading to metric inversion where critical decision-relevant details are missed despite high textual similarity.
To resolve this, recent approaches shift toward behavior-aligned training and reinforcement learning objectives that prioritize functional consistency over literal reconstruction. Methods like Behavior Consistency Reward (BehR) and RLVR-World optimize world models to ensure that predicted states induce the same action distributions as real environments, thereby improving downstream task utility and lookahead planning performance in web agents.
The material examines a train–inference mismatch in web-agent world models. In these systems, a world model is typically trained with supervised next-state prediction: given an observation and an action, it learns to predict the resulting page state or observation. Downstream, a ranker or policy uses those predicted states to score candidate actions and choose the one most likely to succeed. The key insight is that this objective is not aligned with the ranker’s actual need. Predicting a single next state accurately in a reconstruction sense does not guarantee that the model can discriminate well among the states that would result from different candidate actions, especially when those actions produce visually similar, partially overlapping, or only subtly different outcomes.
Its main contribution is to reframe web-agent world modeling as a discriminative rather than purely generative or reconstructive task. Rather than asking only, “What is the next state after this action?”, the work emphasizes the question, “Which candidate action leads to the more desirable or distinguishable next state?” This shifts the design focus toward state representations and prediction objectives that preserve action-relevant differences, improve ranking fidelity, and reduce the gap between world-model training and action-selection performance.
This matters because web agents operate in high-dimensional, noisy environments where small differences in predicted page states can determine whether a click, navigation, or form-filling action is correct. If the world model is optimized for prediction accuracy alone, it may produce states that are plausible but poorly calibrated for ranking, leading to brittle or suboptimal action selection. By aligning world-model learning with discriminative action evaluation, the work points toward more reliable decision-making, better use of limited interaction data, and a clearer path to world models that support planning rather than only imitation.