Identifies limitations of adapting image diffusion models for geometry estimation: independent training loses correlations between targets while joint fine-tuning of modified backbones incurs high cost.

Topological visualization of Video Generative Models as Geometry Learner
Brave API

GeoNeXt addresses these limitations by repurposing pretrained video generative models as a unified, data-efficient framework for geometry estimation. It formulates the task as next-frame prediction, leveraging the temporal coherence and rich priors of video models to jointly model images and geometry targets (depth and surface normals) in a lockstep denoising trajectory. This approach enables effective learning with substantially less training data while outperforming previous task-specific and unified generative competitors.

Generated Aug 31, 2026
Open-Weights Reasoning

The material examines how diffusion-based generative models can be repurposed for geometry estimation, with a focus on a common but limited strategy: adapting image diffusion backbones to predict geometric signals. It argues that existing approaches face a structural trade-off. Training separate image-diffusion models independently for different geometry targets treats each task in isolation and discards the strong correlations among geometric quantities—such as consistency between depth, surface structure, and motion. Conversely, jointly fine-tuning a modified image backbone to predict multiple geometry targets at once can preserve some of those cross-task dependencies, but does so at high computational and engineering cost.

The key insight is that video generative models may be a more natural “geometry learner” than image-only diffusion models. Because video models are trained on spatiotemporally coherent sequences, they already learn dependencies across frames, objects, surfaces, and motion that are directly relevant to inferring scene geometry. Rather than forcing an image prior into a multi-task geometry estimator, the material suggests exploiting the sequential and multi-view structure of video to capture geometric correlations more efficiently.

This matters because dense geometry estimation is a core bottleneck for 3D-aware video understanding, reconstruction, robotics, and generative scene editing. If video-native generative priors can encode geometric consistency with less task-specific fine-tuning than modified image diffusion backbones, they could offer a more scalable and cost-effective route to geometry-aware vision models. The contribution is therefore both diagnostic and directional: it clarifies why current image-diffusion adaptations underperform and motivates video generative models as a promising foundation for learning geometry.

Generated Aug 31, 2026
Sources