V-JEPA: Learning Video Representations by Feature Prediction

((8, 14, 14), 1568) This article follows V-JEPA from its video input to a pretrained encoder and a few small experiments. The complete notebook, including saved outputs and setup notes, is available as a downloadable ZIP, or you can read the interactive SolveIt dialogue.

JEPA stands for Joint-Embedding Predictive Architecture. It is an architecture and training approach for learning representations by predicting one part of an input from another in feature space. I-JEPA applies this idea to images, and V-JEPA extends it to video, where the model can learn from changes across frames. Both predict feature vectors rather than reconstructing pixels.

For the original I-JEPA and V-JEPA, the encoder is the main result of pretraining. It turns an image or video into features that other models can use. A "V-JEPA model" can therefore mean the pretrained encoder produced by this process, as well as the architecture used to train it.