Presents the MindTopo benchmark to evaluate foundation models on five cognitive-science-grounded topological spatial relations that are invariant under deformation.

Topological visualization of MindTopo: Can Foundation Models Reason in Topological Space?
Brave API

MindTopo is a benchmark designed to evaluate whether multimodal large language models (MLLMs) possess topological intuition, assessing their ability to recognize spatial relations that remain invariant under continuous deformation. The benchmark is grounded in Piaget’s classification of spatial understanding and evaluates five core topological properties: continuity, separation, order, enclosure, and knots.

Evaluation occurs across two cognitive levels: reasoning, which tests static visual question-answering over rendered scenes, and planning, which requires interactive agents to manipulate environments while preserving topological constraints. The benchmark consists of 11,030 instances across 13 procedurally generated task types with controllable difficulty, revealing a significant performance gap between static recognition and active planning in current foundation models.

Generated 23d ago
Open-Weights Reasoning

MindTopo introduces a benchmark for testing whether foundation models can reason in topological space rather than merely recognizing low-level visual or metric cues. The central idea is to evaluate models on five cognitive-science-grounded topological spatial relations—qualitative relationships that remain valid under continuous deformation, such as changes in shape, scale, or perspective. By anchoring the tasks in topological structure, the benchmark shifts the focus from appearance-based spatial perception to more abstract relational reasoning about how objects or regions are organized in space.

A key contribution of the work is that it provides a controlled, principled way to probe a capability that is often assumed but rarely measured directly: deformation-invariant spatial reasoning. Because topological relations are preserved under broad transformations, correct reasoning should not depend on precise geometry, pixel layout, or superficial visual features. This makes the benchmark especially useful for diagnosing whether foundation models possess robust relational representations or instead rely on brittle heuristics tied to particular visual configurations. In that sense, MindTopo serves as a stress test for the kind of spatial abstraction that is central to human cognition but difficult to infer from standard visual benchmarks.

The material matters because topological reasoning is foundational to many areas of AI, including multimodal understanding, robotics, navigation, manipulation, and scene structure inference. A model that can reason topologically is better positioned to generalize across deformations, partial observations, and non-rigid scenes. By connecting cognitive-science notions of spatial relations to foundation-model evaluation, MindTopo offers a useful diagnostic for progress toward more robust, human-like spatial intelligence.

Generated 23d ago
Sources