Introduces the MNIST-PRO benchmark that isolates agentic perception by turning digit recognition into a sequential, glimpse-based search task with lookback constraints.

Topological visualization of MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents
Brave API

The search context provided does not contain information regarding MNIST-PRO or its specific features as a partially observable world for AI agents. The available results describe MMMU-Pro, which is a multimodal benchmark for academic disciplines, and the original MNIST dataset, which is a static collection of 70,000 grayscale images used for standard digit recognition tasks. There is no mention of a sequential, glimpse-based search variant of MNIST in the provided sources.

Generated Sep 1, 2026
Open-Weights Reasoning

MNIST-PRO reframes the classic MNIST digit-recognition problem as a partially observable sequential task for AI agents. Instead of presenting a full image and asking for a single classification, the benchmark exposes the digit through a limited sequence of “glimpses,” forcing an agent to decide what to observe, when to stop, and how to infer the target digit from incomplete evidence. The inclusion of lookback constraints further tightens the setting: the agent cannot simply store and reinspect an unbounded history of observations, making memory, query planning, and online inference central to success.

The key contribution is a controlled, minimal environment that isolates agentic perception from the broader stack of language, tool use, or long-horizon planning. By turning recognition into a glimpse-based search problem, MNIST-PRO makes explicit the parts of perception that are often implicit in standard vision benchmarks: active information gathering, uncertainty management, and decision-making under partial observability. It also provides a reproducible testbed for comparing different agent architectures, including reactive classifiers, memory-augmented planners, and policies that explicitly model observation budgets.

This matters because many modern agent systems are evaluated on end-to-end tasks where perception, reasoning, and action are tightly entangled, making it difficult to diagnose where failures originate. MNIST-PRO offers a low-cost, interpretable benchmark for probing whether an agent can acquire useful evidence efficiently rather than merely recognizing fully specified inputs. As a result, it can serve as a useful stress test for active perception, bounded-memory reasoning, and robustness to delayed or incomplete observation before evaluating agents in more complex, high-dimensional environments.

Generated Sep 1, 2026
Sources