Introduces LeVJEPA, the first video encoder trained under a collapse-free LeJEPA objective that avoids architectural asymmetries and pixel-space reconstruction.

Topological visualization of LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics
Brave API

LeVJEPA is the first video encoder trained under LeJEPA’s collapse-free objective, eliminating the need for architectural asymmetries (like target encoders or stop-gradients) and pixel-space reconstruction. It utilizes a single encoder optimized with an invariance loss over global and local views, regularized by SIGReg to prevent representation collapse.

This approach enables two key advantages: FLOP efficiency and block-causal attention. LeVJEPA matches or exceeds V-JEPA 2 accuracy while requiring 5.6 to 20.8 times less pretraining compute. Additionally, it supports causal frame representations where each frame depends only on past observations, facilitating streaming inference and autoregressive modeling without accuracy loss.

Generated Aug 30, 2026
Open-Weights Reasoning

LeVJEPA: Efficient & Scalable Video Pretraining without the Heuristics introduces a video pretraining framework built around a collapse-free latent predictive objective, referred to as LeJEPA. Rather than training a video encoder to reconstruct pixels or to separate positive/negative samples in embedding space, the method learns by predicting latent representations of related video segments or views. Its central contribution is extending this class of joint-embedding predictive learning to video in a way that does not depend on the usual stabilization heuristics, such as architectural asymmetries, stop-gradient pathways, momentum teachers, or pixel-space reconstruction targets. In doing so, it positions video representation learning as a direct latent-space prediction problem instead of a pipeline assembled from ad hoc tricks to prevent collapse.

The broader significance is that this approach targets two common bottlenecks in video self-supervised learning: computational efficiency and architectural complexity. Pixel-space objectives can be bandwidth-heavy and may overemphasize low-level appearance, while contrastive or asymmetric predictive methods often require careful engineering and large-scale batching to remain stable. By using a collapse-free latent objective, LeVJEPA offers a simpler and potentially more scalable recipe for pretraining video encoders. If the method performs competitively on downstream tasks, it would be an important step toward JEPA-style video foundation models that are easier to train, easier to scale, and less dependent on fragile heuristics.

Generated Aug 30, 2026
Sources