{"article":{"slug":"v-jepa-learning-video-representations-by-feature-prediction","title":"V-JEPA: Learning Video Representations by Feature Prediction","subtitle":null,"summary":"A hands-on walkthrough of Meta’s V-JEPA: video tubelets, feature prediction, pretrained encoder features, temporal tests, and a small action-classification experiment.","content_type":"tutorial","language":"en","canonical_url":"https://exploringml.com/posts/how-v-jepa-works/","author":{"name":"David Gwyer","url":"https://exploringml.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"ExploringML","url":"https://exploringml.com/","listing_slug":null,"listing":null},"topics":[{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Tutorial","slug":"tutorial","url":"https://listedarticles.com/topics/tutorial"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":5743,"reading_minutes":25,"published_at":"2026-09-25T09:59:57.000Z","added_at":"2026-09-26T06:12:56.069Z","updated_at":"2026-09-26T06:12:56.069Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/v-jepa-learning-video-representations-by-feature-prediction","markdown_url":"https://listedarticles.com/articles/v-jepa-learning-video-representations-by-feature-prediction.md","example":false,"citation":"David Gwyer, ExploringML. \"V-JEPA: Learning Video Representations by Feature Prediction.\" 25 Sept 2026. https://exploringml.com/posts/how-v-jepa-works/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://exploringml.com/posts/how-v-jepa-works/"},"body_markdown":"# V-JEPA: Learning Video Representations by Feature Prediction\n\n`((8, 14, 14), 1568)`\nThis article follows V-JEPA from its video input to a pretrained encoder and a few small experiments. The complete notebook, including saved outputs and setup notes, is available as a [downloadable ZIP](https://exploringml.com/how-v-jepa-works.zip), or you can read the [interactive SolveIt dialogue](https://share.solveit.pub/d/d59a561edbdff9e2e413b6a7a2060386).\n\nJEPA stands for Joint-Embedding Predictive Architecture. It is an architecture and training approach for learning representations by predicting one part of an input from another in feature space. [I-JEPA](https://arxiv.org/abs/2301.08243) applies this idea to images, and [V-JEPA](https://arxiv.org/abs/2404.08471) extends it to video, where the model can learn from changes across frames. Both predict feature vectors rather than reconstructing pixels.\n\n**For the original I-JEPA and V-JEPA, the encoder is the main result of pretraining.** It turns an image or video into features that other models can use. A \"V-JEPA model\" can therefore mean the pretrained encoder produced by this process, as well as the architecture used to train it.\n\n## One Pretraining Step\n\nTake a video of a hand lifting a cup and hide some patches from the context branch. The context encoder represents the visible content. A smaller predictor uses those features to estimate vectors for the hidden positions. The target encoder sees the full clip and supplies the vectors used to check those predictions.\n\nThe two diagrams presented below follow the forward pass of one pretraining step: computing features, making predictions and measuring their error. Each prediction is compared with the target vector at the same position in space and time. The weight updates that follow this comparison complete the training step.\n\nA tubelet is a video patch that spans consecutive frames. Here, it covers the same 16 × 16-pixel region in two sampled frames. The embedding layer turns each tubelet into a 1,024-value vector, then adds a position vector of the same size.\n\nThe resulting vector represents one token, an item in the transformer's input sequence. It carries information about the tubelet's content and its position in space and time. The context branch keeps only the visible tokens; the transformer then updates each one using information from the others.\n\n```\nflowchart TD\n X[\"Video clip<br/>16 sampled RGB frames<br/>224 × 224 pixels each\"]\n subgraph C[\"CONTEXT ENCODER\"]\n C1[\"Patch embedding layer<br/>Each RGB tubelet:<br/>2 frames × 16 × 16 pixels<br/>Learned projection<br/>1,024 values per tubelet<br/>1,568 vectors per clip\"] --> C2[\"Add position vectors<br/>1,024 + 1,024 → 1,024\"]\n C2 --> C3[\"Keep visible tokens<br/>Remove hidden tokens\"]\n C3 --> C4[\"Transformer blocks<br/>Attention across<br/>visible tokens<br/>Final layer normalisation\"]\n end\n subgraph T[\"TARGET ENCODER\"]\n T1[\"Patch embedding layer<br/>Each RGB tubelet:<br/>2 frames × 16 × 16 pixels<br/>Learned projection<br/>1,024 values per tubelet<br/>1,568 vectors per clip\"] --> T2[\"Add position vectors<br/>1,024 + 1,024 → 1,024\"]\n T2 --> T3[\"Keep all tokens\"]\n T3 --> T4[\"Transformer blocks<br/>Attention across<br/>all tokens<br/>Final layer normalisation\"]\n end\n X ----> C1\n X ----> T1\n C4 --> V[\"Visible features<br/>To predictor below\"]\n T4 --> F[\"Full-clip features<br/>To target selection below\"]\n classDef online fill:#eef4fa,stroke:#7392ad,color:#243447\n classDef teacher fill:#f0f7f5,stroke:#7b9c92,color:#243447\n class C1,C2,C3,C4,V online\n class T1,T2,T3,T4,F teacher\n style C fill:#fafcfe,stroke:#a6b8c8\n style T fill:#f9fcfb,stroke:#a1b9b0\n```\nEach tubelet covers the same 16 × 16-pixel region in two sampled frames. It contains 2 × 16 × 16 × 3 = 1,536 RGB values. The patch embedding layer maps this whole block to one 1,024-value vector using a learned 3D convolution. It reads the values from both frames together; it does not first add or average the two patches. Grouping the pixels into tubelets and projecting them are performed by this layer.\n\nA 224 × 224 frame has 14 × 14 spatial positions. Sixteen sampled frames form eight pairs, giving 8 × 14 × 14 = 1,568 tubelets and therefore 1,568 initial embedding vectors per clip. Each receives a 1,024-value position vector through elementwise addition. The context branch then removes hidden tokens before attention; the target branch retains all of them. [Patch embedding layer](https://github.com/facebookresearch/jepa/blob/main/src/models/utils/patch_embed.py), [encoder](https://github.com/facebookresearch/jepa/blob/main/src/models/vision_transformer.py).\n\nAttention lets each token use information from other tokens in its input sequence. The context transformer can attend across the retained visible tubelets; the target transformer can attend across the whole clip. Each output still corresponds to one tubelet position, but now incorporates context from other positions. Final layer normalisation normalises the values within each output vector while preserving its 1,024-value width and position in the sequence.\n\nThe predictor uses the same kind of fixed sine-and-cosine position encoding as the encoder, generated at width 384. Time, row and column encodings occupy portions of one position vector; they are not three separate 384-value vectors added to a token. [Position encoding implementation](https://github.com/facebookresearch/jepa/blob/main/src/models/utils/pos_embs.py).\n\n```\nflowchart TD\n V[\"Visible features<br/>1,024 values per token\"]\n subgraph P[\"PREDICTOR\"]\n R[\"Linear projection<br/>1,024 → 384\"] --> A[\"For each visible token<br/>Projected feature + position<br/>384 + 384 → 384\"]\n H[\"For each hidden position<br/>Learned mask token<br/>+ position vector<br/>384 + 384 → 384\"]\n A --> J[\"Append placeholders<br/>to visible tokens<br/>One longer sequence<br/>384 values per token\"]\n H --> J\n J --> B[\"12 transformer blocks<br/>Use visible context<br/>to predict hidden features<br/>Final layer normalisation\"]\n B --> O[\"Keep hidden outputs<br/>Project 384 → 1,024\"]\n end\n V ----> R\n O --> Y[\"Predicted vectors<br/>One per hidden position\"]\n F[\"Full-clip target features\"] --> S[\"Normalise features<br/>Select matching<br/>hidden positions\"]\n S --> Z[\"Target vectors<br/>1,024 values each\"]\n Y --> L[\"L1 loss<br/>Match hidden positions\"]\n Z --> L\n L -.-> U[\"Weight updates<br/>Backpropagation:<br/>context encoder + predictor<br/>Target: moving average\"]\n style U fill:#fffaf0,stroke:#b9a587,stroke-dasharray:4 3\n classDef online fill:#eef4fa,stroke:#7392ad,color:#243447\n classDef target fill:#f0f7f5,stroke:#7b9c92,color:#243447\n classDef mask fill:#fff1df,stroke:#c27828,color:#804a16\n class R,A,J,B,O,V,Y online\n class F,S,Z target\n class H mask\n style P fill:#fafcfe,stroke:#a6b8c8\n```\nEach visible feature is projected to 384 values, then added elementwise to its own position vector. Each hidden position instead gets a learned placeholder, called a mask token, plus its own position vector. The placeholder contains no hidden image content. Both operations produce separate 384-value token vectors.\n\nEvery forward pass starts with placeholders made from the current learned mask token and each hidden position's vector. They do not carry guesses over from the previous pass.\n\n\"Append\" joins the two lists of tokens. For example, 10 visible tokens and 6 hidden-position tokens produce a sequence of 16 vectors, each 384 values wide. Attention lets the placeholders use the visible features and interact with other tokens. Their outputs become predictions for the requested hidden positions. Each requested hidden position gets one predicted vector. Sixteen hidden patches would mean sixteen predictions, but the mask determines that count; it is not fixed by the model. [Predictor implementation](https://github.com/facebookresearch/jepa/blob/main/src/models/predictor.py).\n\nThese diagrams use V-JEPA ViT-L/16, with encoder width 1,024. For I-JEPA ViT-H/14, that width is 1,280. Its predictor also projects to its narrower internal width before adding predictor position vectors. [I-JEPA implementation](https://github.com/facebookresearch/ijepa/blob/main/src/models/vision_transformer.py).\n\nThe dashed arrow beneath the loss marks the weight updates after the forward pass. Backpropagation uses the prediction error to update both the predictor and the context encoder. The context encoder learns to produce features that help predict the hidden content, even though it produces outputs only for visible positions.\n\nThe target encoder receives no gradient update. Its weights follow an exponential moving average of the context encoder's weights, giving a gradually changing reference. This comparison and update happen on every training step. The target vectors are used for the loss and are never fed into the predictor. [V-JEPA method](https://arxiv.org/html/2404.08471v1#S3).\n\nIn ViT-L/16, each video patch becomes a vector with 1,024 coordinates. Spatial and temporal position information is added before the encoder's transformer processes the visible tokens.\n\nThe predictor projects the encoder outputs to 384 coordinates and adds position information at that width. Hidden positions receive learned mask tokens with their own position information. Twelve transformer blocks process the visible and hidden-position tokens together. An output projection then maps the hidden-position results back to 1,024 coordinates for comparison with the target vectors. These widths count the coordinates in each vector. The number of patch positions is a separate dimension. [Configuration](https://github.com/facebookresearch/jepa/blob/main/configs/pretrain/vitl16.yaml), [predictor implementation](https://github.com/facebookresearch/jepa/blob/main/src/models/predictor.py).\n\nThe smaller predictor is intended to encourage the encoder to learn representations that make prediction easier. Reduced width alone does not guarantee useful features or prevent collapse, where different inputs produce the same representation.\n\nAfter pretraining, we can use the encoder to extract features for a downstream task, meaning a task we want to solve with the learned representations. For action classification, a classifier learns to map video features to action labels. The encoder can remain frozen, with its weights fixed, or be fine-tuned for that task. Ordinary feature extraction uses the encoder without the masking, predictor or target-comparison branch.\n\nUsing the pretrained encoder for action classification:\n\n```\ngraph LR\n V[\"Video clip\"] --> E[\"Pretrained encoder\"] --> F[\"Video features\"] --> C[\"Trained classifier\"] --> A[\"Action label\"]\n classDef model fill:#eef4fa,stroke:#7392ad,color:#243447\n class E,C model\n```\nOnce the classifier has been trained, this forward pass produces a prediction for a new clip. No weight update is needed for inference.\n\n## What Changes When the Picture Moves?\n\nA photograph shows the hand and cup at one instant. Across a video, their positions change: the hand approaches, grips the cup and lifts it. V-JEPA's input includes these changes over time, while I-JEPA receives a single image.\n\nAn image supplies patches arranged in rows and columns. A video adds a time dimension: each patch covers a fixed region across consecutive sampled frames. Those pixels become embedding vectors, which form the token sequence processed by the transformer.\n\nThe [original V-JEPA method](https://arxiv.org/html/2404.08471v1#S3) predicts hidden regions within a clip. The visible context can include information from across that clip, so this training task should not be treated as evidence of forecasting future events or planning actions.\n\nNow that we understand the architecture, we can see what V-JEPA's representations look like in practice. We'll take a real video clip, select a few frames, and inspect the features produced by the pretrained encoder. Then we'll hide parts of the clip and examine what the model predicts for them.\n\nBecause V-JEPA is designed for video, we'll also change the temporal information and observe how its representations respond. Finally, we'll train a small action classifier on top of frozen encoder features. This won't reproduce a full benchmark, but it will give us a concrete test of what the model has learned from our examples.\n\nThe pretrained weights and GPU-intensive computation will run on Modal, and the results will be returned directly.\n\n## The Model We'll Explore\n\nWe'll use Meta's original V-JEPA ViT-L/16 with 224 × 224-pixel input frames. Its released checkpoint, `vitl16.pth.tar`, is listed in [Meta's model zoo](https://github.com/facebookresearch/jepa#pretrained-models). It is smaller than the ViT-H variants.\n\nThe table gives its input dimensions and encoder feature width, which is the number of coordinates in each output vector.\n\n| Detail | V-JEPA ViT-L/16 | \n|---|---|\n| Sampled input frames | 16 | \n| Frame resolution | 224 × 224 pixels | \n| Spatial patch size | 16 × 16 pixels | \n| Temporal patch depth | 2 sampled frames | \n| Patch grid (time × rows × columns) | 8 × 14 × 14 | \n| Tokens in the complete clip | 1,568 | \n| Encoder feature width | 1,024 | \n\nThese values come from the [ViT-L/16 pretraining configuration](https://github.com/facebookresearch/jepa/blob/main/configs/pretrain/vitl16.yaml) and [vision-transformer implementation](https://github.com/facebookresearch/jepa/blob/main/src/models/vision_transformer.py).\n\nA video patch contains the same 16 × 16-pixel region across two sampled frames. Its location stays fixed in the image while objects can move through it.\n\nThe model converts these pixels into an embedding vector. That vector occupies one position in the transformer's input sequence, where we call it a token. \"Patch\" describes the raw input region, \"embedding\" its numerical representation, and \"token\" its place in the sequence.\n\nThe official configuration uses `sampling_rate: 4`, so consecutive frames in the model's input need not be adjacent in the source video. We first sample frames, then group that sequence into pairs.\n\nWhen we load the checkpoint, we'll check for the context-encoder, target-encoder and predictor weights needed for the hidden-region experiment.\n\n### Count the Video Patches\n\nI-JEPA ViT-H/14 divides a 224 × 224 image into 14 × 14-pixel patches, giving 16 rows × 16 columns = 256 tokens. V-JEPA ViT-L/16 uses larger spatial patches, so each frame pair has 14 rows × 14 columns = 196 positions.\n\nSixteen sampled frames give eight non-overlapping pairs. Multiplying the grid dimensions gives the total token count as shown below. Note: The `//` operator performs whole-number division.\n\nThe output `((8, 14, 14), 1568)` gives the grid in (time, rows, columns) order and the total of 1,568 tokens. It counts patch positions; it does not contain their learned vectors.\n\nNext, we'll select the 16 input frames from an actual video.\n\n## Choosing the Frames\n\nWe'll use a short basketball clip from UCF101, also used in [Hugging Face's video-classification guide](https://huggingface.co/docs/transformers/tasks/video_classification). A player moving towards the basket gives us motion to follow across the sampled frames.\n\nFor this first inspection, we'll take 16 frames at a spacing of four source frames. We'll display them before resizing or normalising the pixels. This is lightweight video preparation in SolveIt; the encoder and GPU work will run on Modal later. OpenCV will decode the video, so the next cell installs the package we need.\n\n`Note: you may need to restart the kernel to use updated packages.``('vjepa_basketball_sample.avi', 545862)`\nThe cell above downloads the clip if needed and reports its filename and file size in bytes. Here's the five-second video. Play it to see the movement, then compare it with the sampled frames below.\n\n`{'frames': 125, 'fps': 25.0, 'size (width, height)': (320, 240)}`\nThe clip has 125 frames at 25 frames per second. We'll choose a window near the middle and keep every fourth frame. That makes neighbouring sampled frames 0.16 seconds apart.\n\nSixteen samples contain 15 gaps, so the first and last samples will be 2.4 seconds apart. This is our simple, repeatable sampling choice for the walkthrough; training can sample different windows.\n\n`[32, 36, 40, 44, 48, 52, 56, 60, 64, 68, 72, 76, 80, 84, 88, 92]`\nFrame numbers start at zero. We'll read the selected frames in order and convert OpenCV's BGR colour order to RGB for display. Each item in `sampled_frames` will still be a complete 240 × 320 image, with three colour channels per pixel.\n\n`(16, (240, 320, 3))`\nRead the grid from left to right, one row at a time. The players move towards the basket as the sequence progresses. These are the source frames, including the black borders and broadcast graphics already present in the clip.\n\nThe pair labels connect this sequence to the model's input: frames 32 and 36 form the first pair, frames 40 and 44 the second, and so on. After spatial preprocessing to 224 × 224, each pair supplies 196 tubelets. Each tubelet covers one fixed 16 × 16 region across the two frames, so a whole frame pair produces 196 tokens, not one.\n\nWe now have the 16 RGB frames. Next we'll prepare their pixel values and dimensions for the pretrained encoder.\n\n## Getting the Clip Ready for the Encoder\n\nOur frames are 240 × 320 pixels. We'll resize the shorter side to 256, preserving the aspect ratio, then take the central 224 × 224 region. Every frame gets the same crop, so preprocessing does not introduce artificial camera movement.\n\nThis follows the single-view evaluation recipe in [Meta's implementation](https://github.com/facebookresearch/jepa/blob/main/evals/video_classification_frozen/utils.py). We'll write the steps out using OpenCV so we can inspect them; small interpolation differences mean this is not a bit-for-bit reproduction of that implementation.\n\n`(16, 224, 224, 3)`\nThe crop removes some of the scene, so it is worth checking what remains before interpreting any model output. The array now has shape `(16, 224, 224, 3)`: frames, height, width, RGB channels.\n\nWe'll divide the pixel values by 255, then subtract a fixed mean and divide by a fixed standard deviation for each colour channel. The constants below come from the same evaluation recipe. They are not statistics calculated from this basketball clip.\n\nFinally, we'll put the axes in the order the encoder expects: `(batch, channels, time, height, width)`. A batch is a group of clips processed together; ours contains just one.\n\n`permute` rearranges the axes, preserving the frame order. `unsqueeze(0)` adds the batch axis. The result is a PyTorch tensor, a multidimensional array ready for the model. These values still represent pixels. The encoder's patch-embedding layer will turn them into tubelet embeddings.\n\n`(torch.Size([1, 3, 16, 224, 224]), torch.float32)`\n## A First Look at the Target Encoder's Learned Features\n\nWe can now pass the complete clip through V-JEPA's pretrained **target encoder**. The target encoder should return a 1,024-value vector at each of the clip's 1,568 tubelet positions. Attention lets each position gather information from across the clip, so these vectors describe more than isolated pixels. They are features, not action labels; recognising \"basketball dunk\" would still require a suitable classifier.\n\nWe'll run this target-encoder-only feature extraction on a Modal GPU and return the results. First we need the Modal Python package and an authenticated connection.\n\n`Modal v1.5.5: connection ready`\nA checkpoint stores learned weights. We'll use `vitl16.pth.tar` from [Meta's model zoo](https://github.com/facebookresearch/jepa#model-zoo). The file contains weights for the context encoder (`encoder`), target encoder (`target_encoder`) and predictor, but this walkthrough selects **only** `checkpoint[\"target_encoder\"]`, following the default in the [evaluation code](https://github.com/facebookresearch/jepa/blob/main/evals/video_classification_frozen/eval.py).\n\nDuring V-JEPA pretraining, the target encoder is the moving-average encoder that supplies representation targets. Here, training is over: we use that target encoder by itself as a frozen feature extractor. The context encoder and predictor remain in the checkpoint but are neither instantiated nor run in this section.\n\nThe following code describes the remote software environment. Modal caches this environment, including the downloaded checkpoint, so subsequent runs can reuse it.\n\nBefore running the clip, we'll require every **target-encoder** weight to match the `vit_large` model definition. `strict=True` makes loading fail if any selected target-encoder weights are missing or unexpected, instead of silently leaving part of that encoder randomly initialised.\n\nWe'll also return the checkpoint's top-level keys to verify which components the file contains for later experiments. Seeing `encoder`, `target_encoder` and `predictor` in that list does not mean all three are used here: the loader extracts only `checkpoint[\"target_encoder\"]`.\n\n`eval()` selects evaluation behaviour, while `inference_mode()` avoids building the gradient information needed for training. We pass the complete clip directly to the frozen target encoder, without a mask. It produces one feature vector at every tubelet position, and no weights are updated.\n\nNothing in this call uses V-JEPA's context encoder or predictor. Those components are central to the pretraining task, where visible context is encoded and missing-region representations are predicted, but this is inference after pretraining, using only the learned target encoder to describe the clip.\n\nThe first call also builds the remote environment and downloads the checkpoint, so it will take longer than a cached run. This uses one A10G GPU on Modal for the inference call.\n\n```\nCheckpoint contains: encoder, target_encoder, predictor\nThis run used: target_encoder\nEncoded clip: (1, 1568, 1024)\nJEPA revision: 51c59d51\n```\nThe returned shape is `(1, 1568, 1024)`: one clip, 1,568 target-encoder positions and 1,024 values per position. Strict loading passed for the selected target-encoder weights.\n\nThe checkpoint also contains `encoder` and `predictor` weights, but their presence should not be confused with execution: this run loaded and called only `target_encoder`. No context-encoder or predictor outputs contribute to `features`.\n\nWe can arrange the 1,568 positions back into the familiar grid of eight frame pairs, 14 rows and 14 columns. That only changes how we index the target encoder's output. Each vector already includes information gathered from across the clip through the target encoder's attention layers.\n\n```\n((8, 14, 14, 1024),\n array([ 0.025, -0.954, -0.669,  1.332,  0.463, -0.93 , -0.63 ,  0.265,\n         0.721, -0.162, -0.434,  0.407], dtype=float32))\n```\nThose are the first 12 values of one 1,024-value **target-encoder** vector. Row and column indices start at zero. Individual coordinates do not come with labels such as \"ball\" or \"player\", so this list alone tells us little about what the target encoder has learned. Comparing its vectors across positions or under controlled changes to the clip will be more informative.\n\nThese saved features are target-encoder representations only; they contain no separately computed context-encoder representation and no predictor output. We can save the complete features and the source-code revision here, so inspecting this result again does not require another GPU run.\n\n## What Changes When Time Runs Differently?\n\nThe previous experiment used spatial masking while leaving the clip's timeline intact. We can now hold the 16 sampled images fixed and change only their order. If V-JEPA represented a clip as an unordered collection of pictures, these changes would make no difference. Its tubelets and temporal position vectors give us reason to expect otherwise.\n\nWe'll make two altered clips. One swaps the two frames inside every tubelet pair. The other reverses the complete sequence. Both contain exactly the same normalised frames as the original, with no new pixels and none removed.\n\n`{'original': [32, 36, 40, 44, 48, 52, 56, 60, 64, 68, 72, 76, 80, 84, 88, 92], 'reversed': [92, 88, 84, 80, 76, 72, 68, 64, 60, 56, 52, 48, 44, 40, 36, 32], 'within-pair swap': [36, 32, 44, 40, 52, 48, 60, 56, 68, 64, 76, 72, 84, 80, 92, 88]}`\nThe within-pair version keeps each pair in the same temporal slot but reverses its two frames: `(32, 36)` becomes `(36, 32)`. This tests whether a two-frame tubelet responds to local direction.\n\nThe fully reversed version also moves the pairs to different temporal slots. When we compare its outputs with the original, we'll reverse the eight output groups back into source-pair order. Matching positions will then refer to the same two source frames and spatial patch, although the model encountered them in the opposite temporal direction.\n\n`(torch.Size([1, 3, 16, 224, 224]), torch.Size([1, 3, 16, 224, 224]))`\nWe'll use the same frozen target encoder and the same preprocessing as before. Each clip is encoded separately, so the GPU does not mix information between variants. The output for each one remains `(1, 1568, 1024)`.\n\n`(3, 1568, 1024)`\nCosine similarity compares the direction of two feature vectors while ignoring their overall length. A value of 1 means identical direction, 0 means no directional alignment, and negative values point in opposing directions.\n\nWe'll calculate it for matching spatial positions and source-frame pairs, then average over the 196 positions in each pair. For the fully reversed clip, this requires reversing the eight output groups before comparison.\n\n`(0.6825879216194153, 0.4968155026435852)`\nSwapping the two frames inside every tubelet pair gives a mean cosine similarity of about 0.683. The source images, pair membership and temporal slots are unchanged, so the drop from 1 shows that frame order inside a tubelet affects the encoder's features.\n\nReversing the complete clip lowers the aligned mean similarity to about 0.497. That change includes the reversal inside each pair, different temporal positions for the pairs, and attention across the reversed sequence. This experiment shows sensitivity to temporal order on one clip. It does not establish that the encoder recognises the action, understands cause and effect, or would respond the same way across a dataset.\n\nThe stronger response to complete reversal is consistent with a representation that uses information over time as well as appearance. It is not possible to assign that extra change to one mechanism from this comparison alone, because temporal position and cross-token attention change together.\n\nThe next section will ask a more practical question: can a small classifier use frozen V-JEPA features to separate actions without changing the encoder?\n\n`Saved vjepa_temporal_order.npz`\nAs with the earlier feature-extraction and hidden-region experiments, we save these results locally in SolveIt as a compressed NumPy archive. This avoids repeating the GPU inference when we return to the temporal-order comparison.\n\nThis archive records the target-encoder features for the original, within-pair-swapped and fully reversed clips; the corresponding source-frame orders; the position-wise cosine similarities; and the JEPA source revision used for the run. Keeping the orders and revision beside the arrays makes the comparison easier to interpret and reproduce.\n\nIt can be reopened with `np.load(\"vjepa_temporal_order.npz\")`. Its contents are then available by name, including `data[\"features\"]`, `data[\"pair_swap_cosine\"]` and `data[\"reverse_cosine\"]`.\n\n## Can Frozen V-JEPA Features Separate Actions?\n\nPretraining tells us how V-JEPA learns, but the practical question is what its encoder is useful for afterwards. We'll test that with a **linear probe**: freeze the released target encoder, reduce each clip to one feature vector, and train only a linear classifier to distinguish three actions.\n\nThe experiment has four stages:\n\n1. choose clips from `Archery` ,`BabyCrawling` , and`BasketballDunk` ;\n2. preprocess each clip into the same 16-frame input used above;\n3. average the encoder's 1,568 output tokens into one 1,024-value clip representation; and\n4. fit a three-way linear classifier on those frozen representations.\n\nA linear probe is deliberately limited. If it separates the actions, the class information was already accessible in the pretrained features; the classifier did not teach the encoder a new representation. We use the target encoder here for consistency with the preceding feature experiments and Meta's frozen-evaluation default.\n\nThis is a small demonstration rather than a UCF101 benchmark: only 18 clips are used, from three visually distinct classes. To make the test more meaningful, we split by UCF101 **recording group** rather than by individual clip. Four groups per class are used for training and two different groups for testing, so near-duplicate clips from the same original recording cannot appear on both sides.\n\nThe probe reuses the Modal app, checkpoint image, and strict target-encoder loader from the earlier experiments. The only additional remote dependency is OpenCV, which decodes the UCF101 videos. The next cell adds OpenCV to that environment; the following check spells out the expected clip shape and resulting tubelet count.\n\n`((3, 16, 224, 224), (2, 16, 16), 1568)`\n### Selecting Clips and Splitting by Recording Group\n\nThe downloaded UCF101 subset contains several clips from each recording group. We keep one clip per group, then assign the first four groups in each class to training and the next two to testing. The important unit of separation is the group, not the filename: clips cut from one recording stay on only one side of the split.\n\nBefore preprocessing the clips, we can inspect the exact sample used by the probe. Each row below is one selected recording-group clip; the five columns show evenly spaced source frames from beginning to end. The row labels identify the action and whether the clip belongs to the training or held-out split.\n\nThis is a useful visual check for near-duplicate scenes, uninformative opening or closing frames, and background cues that might make the small classification task easier than the action itself.\n\n### Turning Each Video into One Model Input\n\nThe source clips vary in length, so we choose 16 evenly spaced frames from the beginning to the end of each clip. Each frame then follows the same resize, centre-crop and channel-normalisation recipe used earlier.\n\nThis sampling choice is intentionally simple and deterministic. It gives every video the required shape, but it is not the multi-view evaluation protocol used for a full benchmark.\n\n### Extracting Frozen Clip Representations\n\nFor each selected video, the target encoder produces 1,568 vectors. We average them across space and time to obtain one 1,024-value vector for the complete clip. This pooling discards the location of individual tubelets, but leaves a compact representation suitable for a small classifier.\n\nThe encoder stays in evaluation mode and no gradients are computed. The action labels are collected alongside the features, but they are never supplied to V-JEPA.\n\nWith the selection and preprocessing functions in place, we can run feature extraction once on the GPU. The returned arrays contain the pooled features, labels, split assignments and filenames needed for the local probe; the classifier itself will run in SolveIt rather than on the GPU.\n\n```\n((18, 1024),\n np.int64(12),\n np.int64(6),\n ['Archery', 'BabyCrawling', 'BasketballDunk'])\n```\n### Training the Linear Probe\n\nThe extractor returns 18 frozen clip vectors, each 1,024 values wide. For every class, four UCF101 recording groups supply the training examples and two different groups supply the test examples. That gives us 12 training clips and 6 held-out clips, with no recording group shared across the split.\n\nThe classifier is deliberately small. We normalise every clip vector to unit length, append a constant bias feature, and fit a regularised linear map from the 1,024 encoder features to three class scores. The highest score becomes the predicted action.\n\nThere are only 12 training examples but more than 1,000 feature coordinates, so regularisation matters: without it, many solutions could memorise the training set. The expression below is ridge regression written in its dual form, which solves a 12 × 12 system instead of a 1,025 × 1,025 one. Only these classifier weights learn from the labels, while V-JEPA remains frozen.\n\n```\n(np.float64(1.0),\n np.float64(1.0),\n [(np.int64(0), np.int64(0)),\n  (np.int64(0), np.int64(0)),\n  (np.int64(1), np.int64(1)),\n  (np.int64(1), np.int64(1)),\n  (np.int64(2), np.int64(2)),\n  (np.int64(2), np.int64(2))])\n```\nThe feature matrix first converts to 64-bit floating point for a stable linear solve. Each 1,024-value clip vector is divided by its length, so the classifier compares the direction of the representation rather than allowing vectors with larger magnitudes to dominate. Appending a constant value of one gives the linear classifier a bias term.\n\nThe training labels are converted to one-hot target vectors: for example, an Archery clip has target `[1, 0, 0]`. Ridge regression then learns a regularised linear mapping from the frozen V-JEPA representations to three class scores. The regularisation strength is `0.1`; it discourages excessively large classifier weights in this setting with many more feature coordinates than training examples.\n\nFor each clip, `argmax` selects the class with the highest score. Both reported accuracies are `1.0`, meaning that the classifier correctly labels all 12 training clips and all 6 held-out clips. The `(expected, predicted)` pairs confirm that every held-out numeric label matches its prediction: class `0` is Archery, `1` is BabyCrawling and `2` is BasketballDunk.\n\nPerfect training accuracy is not surprising with so few examples. The held-out result is more informative because its clips come from different recording groups, but six test clips are far too few to estimate performance on UCF101 generally.\n\n```\nv_Archery_g05_c04.avi                  Archery        -> Archery\nv_Archery_g06_c01.avi                  Archery        -> Archery\nv_BabyCrawling_g05_c01.avi             BabyCrawling   -> BabyCrawling\nv_BabyCrawling_g06_c04.avi             BabyCrawling   -> BabyCrawling\nv_BasketballDunk_g05_c01.avi           BasketballDunk -> BasketballDunk\nv_BasketballDunk_g06_c01.avi           BasketballDunk -> BasketballDunk\nHeld-out accuracy: 100% (chance: 33%)\n```\nThe printed rows translate the numeric predictions back into class names. Each row shows the held-out filename, its expected action and the classifier's prediction. All six predictions are correct: two Archery clips, two BabyCrawling clips and two BasketballDunk clips.\n\nThe resulting held-out accuracy is therefore $6/6 = 100\\%$, compared with a 33% chance level for three balanced classes. This demonstrates that these three actions are linearly separable in this small sample of frozen V-JEPA features. It does not establish 100% accuracy on new videos or on the full 101-class dataset; the actions are visually distinct and the test set is intentionally small.\n\nFinally, `vjepa_linear_probe.npz` saves the frozen features, labels, split assignments, filenames, learned classifier weights and class names. This lets us inspect or reuse the probe without downloading the videos and running the encoder again.\n\n### Reading the Held-Out Result\n\nAll six held-out predictions are correct. The confusion matrix below shows their distribution across the three classes.\n\nRows are the true classes, columns are the predicted classes, and each cell counts clips. A perfect result places all six clips on the diagonal.\n\nThe linear probe correctly classifies all six held-out clips, compared with a random 33% chance level for three classes. Its input is only the frozen 1,024-value clip vector. The V-JEPA encoder is never updated, so the result gives us direct evidence that its representation already separates these actions in this small sample.\n\nThe sample is far too small for a performance claim. Archery, baby crawling and basketball dunking are visually distinct, and six test clips cannot represent the variation in UCF101. **The useful point is narrower: a shallow supervised head can read action information from features learned without action labels.**\n\nThe confusion matrix is completely diagonal, matching the six printed predictions. More data and more classes would be needed to test how robust this separation is, but the experiment answers our immediate question: a simple linear classifier can distinguish these three actions from frozen V-JEPA features in this small held-out sample. The encoder itself was not retrained.\n\n## Summary\n\nWe began with raw video: 16 sampled frames from a basketball clip. V-JEPA divided them into 1,568 space-time tubelets and represented each one with a vector of 1,024 learned features. Unlike a pixel-reconstruction model, V-JEPA was trained to predict these representations for hidden parts of a video from the parts it could see.\n\nOur experiments examined that idea from three angles. First, the pretrained predictor estimated the target encoder's features for a withheld central region more accurately than a zero-vector baseline. Changing the hidden pixels had no effect on those predictions, confirming that their contents had not leaked into the visible context.\n\nNext, we changed the order of the same 16 frames. Swapping frames within each tubelet pair changed the encoder's features, and reversing the complete clip changed them more strongly. On this example, V-JEPA's representation therefore depended on temporal order rather than treating the video as an unordered collection of images.\n\nFinally, we froze the encoder and trained only a simple linear classifier on its clip-level features. It correctly separated archery, baby crawling and basketball dunk in a small test set drawn from held-out recording groups. This does not constitute a UCF101 benchmark, but it shows that useful action information was already accessible in the pretrained representation.\n\nTogether, these results illustrate the main idea behind V-JEPA: learning useful video representations by predicting in feature space rather than reconstructing pixels. The experiments are deliberately small: one mask, one temporal-order example and three action classes, but they make the model's data flow and capabilities concrete.\n\nFor the formal method and full evaluations, see the [V-JEPA paper](https://arxiv.org/abs/2404.08471) and the [official JEPA repository](https://github.com/facebookresearch/jepa).\n\nCommunity\n\n## Explore together.\n\nShare what you're learning, ask questions, and swap ideas.\n\n Join the ExploringML Discord.\n\n[Join the community](https://discord.gg/wXhHXc9puA)\n","body_html":"<h1 id=\"v-jepa-learning-video-representations-by-feature-prediction\">V-JEPA: Learning Video Representations by Feature Prediction</h1>\n<p><code>((8, 14, 14), 1568)</code>\nThis article follows V-JEPA from its video input to a pretrained encoder and a few small experiments. The complete notebook, including saved outputs and setup notes, is available as a <a href=\"https://exploringml.com/how-v-jepa-works.zip\" rel=\"nofollow ugc noopener\">downloadable ZIP</a>, or you can read the <a href=\"https://share.solveit.pub/d/d59a561edbdff9e2e413b6a7a2060386\" rel=\"nofollow ugc noopener\">interactive SolveIt dialogue</a>.</p>\n<p>JEPA stands for Joint-Embedding Predictive Architecture. It is an architecture and training approach for learning representations by predicting one part of an input from another in feature space. <a href=\"https://arxiv.org/abs/2301.08243\" rel=\"nofollow ugc noopener\">I-JEPA</a> applies this idea to images, and <a href=\"https://arxiv.org/abs/2404.08471\" rel=\"nofollow ugc noopener\">V-JEPA</a> extends it to video, where the model can learn from changes across frames. Both predict feature vectors rather than reconstructing pixels.</p>\n<p><strong>For the original I-JEPA and V-JEPA, the encoder is the main result of pretraining.</strong> It turns an image or video into features that other models can use. A &quot;V-JEPA model&quot; can therefore mean the pretrained encoder produced by this process, as well as the architecture used to train it.</p>\n<h2 id=\"one-pretraining-step\">One Pretraining Step</h2>\n<p>Take a video of a hand lifting a cup and hide some patches from the context branch. The context encoder represents the visible content. A smaller predictor uses those features to estimate vectors for the hidden positions. The target encoder sees the full clip and supplies the vectors used to check those predictions.</p>\n<p>The two diagrams presented below follow the forward pass of one pretraining step: computing features, making predictions and measuring their error. Each prediction is compared with the target vector at the same position in space and time. The weight updates that follow this comparison complete the training step.</p>\n<p>A tubelet is a video patch that spans consecutive frames. Here, it covers the same 16 × 16-pixel region in two sampled frames. The embedding layer turns each tubelet into a 1,024-value vector, then adds a position vector of the same size.</p>\n<p>The resulting vector represents one token, an item in the transformer&#39;s input sequence. It carries information about the tubelet&#39;s content and its position in space and time. The context branch keeps only the visible tokens; the transformer then updates each one using information from the others.</p>\n<pre><code>flowchart TD\n X[&quot;Video clip&lt;br/&gt;16 sampled RGB frames&lt;br/&gt;224 × 224 pixels each&quot;]\n subgraph C[&quot;CONTEXT ENCODER&quot;]\n C1[&quot;Patch embedding layer&lt;br/&gt;Each RGB tubelet:&lt;br/&gt;2 frames × 16 × 16 pixels&lt;br/&gt;Learned projection&lt;br/&gt;1,024 values per tubelet&lt;br/&gt;1,568 vectors per clip&quot;] --&gt; C2[&quot;Add position vectors&lt;br/&gt;1,024 + 1,024 → 1,024&quot;]\n C2 --&gt; C3[&quot;Keep visible tokens&lt;br/&gt;Remove hidden tokens&quot;]\n C3 --&gt; C4[&quot;Transformer blocks&lt;br/&gt;Attention across&lt;br/&gt;visible tokens&lt;br/&gt;Final layer normalisation&quot;]\n end\n subgraph T[&quot;TARGET ENCODER&quot;]\n T1[&quot;Patch embedding layer&lt;br/&gt;Each RGB tubelet:&lt;br/&gt;2 frames × 16 × 16 pixels&lt;br/&gt;Learned projection&lt;br/&gt;1,024 values per tubelet&lt;br/&gt;1,568 vectors per clip&quot;] --&gt; T2[&quot;Add position vectors&lt;br/&gt;1,024 + 1,024 → 1,024&quot;]\n T2 --&gt; T3[&quot;Keep all tokens&quot;]\n T3 --&gt; T4[&quot;Transformer blocks&lt;br/&gt;Attention across&lt;br/&gt;all tokens&lt;br/&gt;Final layer normalisation&quot;]\n end\n X ----&gt; C1\n X ----&gt; T1\n C4 --&gt; V[&quot;Visible features&lt;br/&gt;To predictor below&quot;]\n T4 --&gt; F[&quot;Full-clip features&lt;br/&gt;To target selection below&quot;]\n classDef online fill:#eef4fa,stroke:#7392ad,color:#243447\n classDef teacher fill:#f0f7f5,stroke:#7b9c92,color:#243447\n class C1,C2,C3,C4,V online\n class T1,T2,T3,T4,F teacher\n style C fill:#fafcfe,stroke:#a6b8c8\n style T fill:#f9fcfb,stroke:#a1b9b0</code></pre>\n<p>Each tubelet covers the same 16 × 16-pixel region in two sampled frames. It contains 2 × 16 × 16 × 3 = 1,536 RGB values. The patch embedding layer maps this whole block to one 1,024-value vector using a learned 3D convolution. It reads the values from both frames together; it does not first add or average the two patches. Grouping the pixels into tubelets and projecting them are performed by this layer.</p>\n<p>A 224 × 224 frame has 14 × 14 spatial positions. Sixteen sampled frames form eight pairs, giving 8 × 14 × 14 = 1,568 tubelets and therefore 1,568 initial embedding vectors per clip. Each receives a 1,024-value position vector through elementwise addition. The context branch then removes hidden tokens before attention; the target branch retains all of them. <a href=\"https://github.com/facebookresearch/jepa/blob/main/src/models/utils/patch_embed.py\" rel=\"nofollow ugc noopener\">Patch embedding layer</a>, <a href=\"https://github.com/facebookresearch/jepa/blob/main/src/models/vision_transformer.py\" rel=\"nofollow ugc noopener\">encoder</a>.</p>\n<p>Attention lets each token use information from other tokens in its input sequence. The context transformer can attend across the retained visible tubelets; the target transformer can attend across the whole clip. Each output still corresponds to one tubelet position, but now incorporates context from other positions. Final layer normalisation normalises the values within each output vector while preserving its 1,024-value width and position in the sequence.</p>\n<p>The predictor uses the same kind of fixed sine-and-cosine position encoding as the encoder, generated at width 384. Time, row and column encodings occupy portions of one position vector; they are not three separate 384-value vectors added to a token. <a href=\"https://github.com/facebookresearch/jepa/blob/main/src/models/utils/pos_embs.py\" rel=\"nofollow ugc noopener\">Position encoding implementation</a>.</p>\n<pre><code>flowchart TD\n V[&quot;Visible features&lt;br/&gt;1,024 values per token&quot;]\n subgraph P[&quot;PREDICTOR&quot;]\n R[&quot;Linear projection&lt;br/&gt;1,024 → 384&quot;] --&gt; A[&quot;For each visible token&lt;br/&gt;Projected feature + position&lt;br/&gt;384 + 384 → 384&quot;]\n H[&quot;For each hidden position&lt;br/&gt;Learned mask token&lt;br/&gt;+ position vector&lt;br/&gt;384 + 384 → 384&quot;]\n A --&gt; J[&quot;Append placeholders&lt;br/&gt;to visible tokens&lt;br/&gt;One longer sequence&lt;br/&gt;384 values per token&quot;]\n H --&gt; J\n J --&gt; B[&quot;12 transformer blocks&lt;br/&gt;Use visible context&lt;br/&gt;to predict hidden features&lt;br/&gt;Final layer normalisation&quot;]\n B --&gt; O[&quot;Keep hidden outputs&lt;br/&gt;Project 384 → 1,024&quot;]\n end\n V ----&gt; R\n O --&gt; Y[&quot;Predicted vectors&lt;br/&gt;One per hidden position&quot;]\n F[&quot;Full-clip target features&quot;] --&gt; S[&quot;Normalise features&lt;br/&gt;Select matching&lt;br/&gt;hidden positions&quot;]\n S --&gt; Z[&quot;Target vectors&lt;br/&gt;1,024 values each&quot;]\n Y --&gt; L[&quot;L1 loss&lt;br/&gt;Match hidden positions&quot;]\n Z --&gt; L\n L -.-&gt; U[&quot;Weight updates&lt;br/&gt;Backpropagation:&lt;br/&gt;context encoder + predictor&lt;br/&gt;Target: moving average&quot;]\n style U fill:#fffaf0,stroke:#b9a587,stroke-dasharray:4 3\n classDef online fill:#eef4fa,stroke:#7392ad,color:#243447\n classDef target fill:#f0f7f5,stroke:#7b9c92,color:#243447\n classDef mask fill:#fff1df,stroke:#c27828,color:#804a16\n class R,A,J,B,O,V,Y online\n class F,S,Z target\n class H mask\n style P fill:#fafcfe,stroke:#a6b8c8</code></pre>\n<p>Each visible feature is projected to 384 values, then added elementwise to its own position vector. Each hidden position instead gets a learned placeholder, called a mask token, plus its own position vector. The placeholder contains no hidden image content. Both operations produce separate 384-value token vectors.</p>\n<p>Every forward pass starts with placeholders made from the current learned mask token and each hidden position&#39;s vector. They do not carry guesses over from the previous pass.</p>\n<p>&quot;Append&quot; joins the two lists of tokens. For example, 10 visible tokens and 6 hidden-position tokens produce a sequence of 16 vectors, each 384 values wide. Attention lets the placeholders use the visible features and interact with other tokens. Their outputs become predictions for the requested hidden positions. Each requested hidden position gets one predicted vector. Sixteen hidden patches would mean sixteen predictions, but the mask determines that count; it is not fixed by the model. <a href=\"https://github.com/facebookresearch/jepa/blob/main/src/models/predictor.py\" rel=\"nofollow ugc noopener\">Predictor implementation</a>.</p>\n<p>These diagrams use V-JEPA ViT-L/16, with encoder width 1,024. For I-JEPA ViT-H/14, that width is 1,280. Its predictor also projects to its narrower internal width before adding predictor position vectors. <a href=\"https://github.com/facebookresearch/ijepa/blob/main/src/models/vision_transformer.py\" rel=\"nofollow ugc noopener\">I-JEPA implementation</a>.</p>\n<p>The dashed arrow beneath the loss marks the weight updates after the forward pass. Backpropagation uses the prediction error to update both the predictor and the context encoder. The context encoder learns to produce features that help predict the hidden content, even though it produces outputs only for visible positions.</p>\n<p>The target encoder receives no gradient update. Its weights follow an exponential moving average of the context encoder&#39;s weights, giving a gradually changing reference. This comparison and update happen on every training step. The target vectors are used for the loss and are never fed into the predictor. <a href=\"https://arxiv.org/html/2404.08471v1#S3\" rel=\"nofollow ugc noopener\">V-JEPA method</a>.</p>\n<p>In ViT-L/16, each video patch becomes a vector with 1,024 coordinates. Spatial and temporal position information is added before the encoder&#39;s transformer processes the visible tokens.</p>\n<p>The predictor projects the encoder outputs to 384 coordinates and adds position information at that width. Hidden positions receive learned mask tokens with their own position information. Twelve transformer blocks process the visible and hidden-position tokens together. An output projection then maps the hidden-position results back to 1,024 coordinates for comparison with the target vectors. These widths count the coordinates in each vector. The number of patch positions is a separate dimension. <a href=\"https://github.com/facebookresearch/jepa/blob/main/configs/pretrain/vitl16.yaml\" rel=\"nofollow ugc noopener\">Configuration</a>, <a href=\"https://github.com/facebookresearch/jepa/blob/main/src/models/predictor.py\" rel=\"nofollow ugc noopener\">predictor implementation</a>.</p>\n<p>The smaller predictor is intended to encourage the encoder to learn representations that make prediction easier. Reduced width alone does not guarantee useful features or prevent collapse, where different inputs produce the same representation.</p>\n<p>After pretraining, we can use the encoder to extract features for a downstream task, meaning a task we want to solve with the learned representations. For action classification, a classifier learns to map video features to action labels. The encoder can remain frozen, with its weights fixed, or be fine-tuned for that task. Ordinary feature extraction uses the encoder without the masking, predictor or target-comparison branch.</p>\n<p>Using the pretrained encoder for action classification:</p>\n<pre><code>graph LR\n V[&quot;Video clip&quot;] --&gt; E[&quot;Pretrained encoder&quot;] --&gt; F[&quot;Video features&quot;] --&gt; C[&quot;Trained classifier&quot;] --&gt; A[&quot;Action label&quot;]\n classDef model fill:#eef4fa,stroke:#7392ad,color:#243447\n class E,C model</code></pre>\n<p>Once the classifier has been trained, this forward pass produces a prediction for a new clip. No weight update is needed for inference.</p>\n<h2 id=\"what-changes-when-the-picture-moves\">What Changes When the Picture Moves?</h2>\n<p>A photograph shows the hand and cup at one instant. Across a video, their positions change: the hand approaches, grips the cup and lifts it. V-JEPA&#39;s input includes these changes over time, while I-JEPA receives a single image.</p>\n<p>An image supplies patches arranged in rows and columns. A video adds a time dimension: each patch covers a fixed region across consecutive sampled frames. Those pixels become embedding vectors, which form the token sequence processed by the transformer.</p>\n<p>The <a href=\"https://arxiv.org/html/2404.08471v1#S3\" rel=\"nofollow ugc noopener\">original V-JEPA method</a> predicts hidden regions within a clip. The visible context can include information from across that clip, so this training task should not be treated as evidence of forecasting future events or planning actions.</p>\n<p>Now that we understand the architecture, we can see what V-JEPA&#39;s representations look like in practice. We&#39;ll take a real video clip, select a few frames, and inspect the features produced by the pretrained encoder. Then we&#39;ll hide parts of the clip and examine what the model predicts for them.</p>\n<p>Because V-JEPA is designed for video, we&#39;ll also change the temporal information and observe how its representations respond. Finally, we&#39;ll train a small action classifier on top of frozen encoder features. This won&#39;t reproduce a full benchmark, but it will give us a concrete test of what the model has learned from our examples.</p>\n<p>The pretrained weights and GPU-intensive computation will run on Modal, and the results will be returned directly.</p>\n<h2 id=\"the-model-we-ll-explore\">The Model We&#39;ll Explore</h2>\n<p>We&#39;ll use Meta&#39;s original V-JEPA ViT-L/16 with 224 × 224-pixel input frames. Its released checkpoint, <code>vitl16.pth.tar</code>, is listed in <a href=\"https://github.com/facebookresearch/jepa#pretrained-models\" rel=\"nofollow ugc noopener\">Meta&#39;s model zoo</a>. It is smaller than the ViT-H variants.</p>\n<p>The table gives its input dimensions and encoder feature width, which is the number of coordinates in each output vector.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Detail</th><th>V-JEPA ViT-L/16</th></tr></thead><tbody><tr><td>Sampled input frames</td><td>16</td></tr><tr><td>Frame resolution</td><td>224 × 224 pixels</td></tr><tr><td>Spatial patch size</td><td>16 × 16 pixels</td></tr><tr><td>Temporal patch depth</td><td>2 sampled frames</td></tr><tr><td>Patch grid (time × rows × columns)</td><td>8 × 14 × 14</td></tr><tr><td>Tokens in the complete clip</td><td>1,568</td></tr><tr><td>Encoder feature width</td><td>1,024</td></tr></tbody></table></div>\n<p>These values come from the <a href=\"https://github.com/facebookresearch/jepa/blob/main/configs/pretrain/vitl16.yaml\" rel=\"nofollow ugc noopener\">ViT-L/16 pretraining configuration</a> and <a href=\"https://github.com/facebookresearch/jepa/blob/main/src/models/vision_transformer.py\" rel=\"nofollow ugc noopener\">vision-transformer implementation</a>.</p>\n<p>A video patch contains the same 16 × 16-pixel region across two sampled frames. Its location stays fixed in the image while objects can move through it.</p>\n<p>The model converts these pixels into an embedding vector. That vector occupies one position in the transformer&#39;s input sequence, where we call it a token. &quot;Patch&quot; describes the raw input region, &quot;embedding&quot; its numerical representation, and &quot;token&quot; its place in the sequence.</p>\n<p>The official configuration uses <code>sampling_rate: 4</code>, so consecutive frames in the model&#39;s input need not be adjacent in the source video. We first sample frames, then group that sequence into pairs.</p>\n<p>When we load the checkpoint, we&#39;ll check for the context-encoder, target-encoder and predictor weights needed for the hidden-region experiment.</p>\n<h3 id=\"count-the-video-patches\">Count the Video Patches</h3>\n<p>I-JEPA ViT-H/14 divides a 224 × 224 image into 14 × 14-pixel patches, giving 16 rows × 16 columns = 256 tokens. V-JEPA ViT-L/16 uses larger spatial patches, so each frame pair has 14 rows × 14 columns = 196 positions.</p>\n<p>Sixteen sampled frames give eight non-overlapping pairs. Multiplying the grid dimensions gives the total token count as shown below. Note: The <code>//</code> operator performs whole-number division.</p>\n<p>The output <code>((8, 14, 14), 1568)</code> gives the grid in (time, rows, columns) order and the total of 1,568 tokens. It counts patch positions; it does not contain their learned vectors.</p>\n<p>Next, we&#39;ll select the 16 input frames from an actual video.</p>\n<h2 id=\"choosing-the-frames\">Choosing the Frames</h2>\n<p>We&#39;ll use a short basketball clip from UCF101, also used in <a href=\"https://huggingface.co/docs/transformers/tasks/video_classification\" rel=\"nofollow ugc noopener\">Hugging Face&#39;s video-classification guide</a>. A player moving towards the basket gives us motion to follow across the sampled frames.</p>\n<p>For this first inspection, we&#39;ll take 16 frames at a spacing of four source frames. We&#39;ll display them before resizing or normalising the pixels. This is lightweight video preparation in SolveIt; the encoder and GPU work will run on Modal later. OpenCV will decode the video, so the next cell installs the package we need.</p>\n<p><code>Note: you may need to restart the kernel to use updated packages.</code><code>(&#39;vjepa_basketball_sample.avi&#39;, 545862)</code>\nThe cell above downloads the clip if needed and reports its filename and file size in bytes. Here&#39;s the five-second video. Play it to see the movement, then compare it with the sampled frames below.</p>\n<p><code>{&#39;frames&#39;: 125, &#39;fps&#39;: 25.0, &#39;size (width, height)&#39;: (320, 240)}</code>\nThe clip has 125 frames at 25 frames per second. We&#39;ll choose a window near the middle and keep every fourth frame. That makes neighbouring sampled frames 0.16 seconds apart.</p>\n<p>Sixteen samples contain 15 gaps, so the first and last samples will be 2.4 seconds apart. This is our simple, repeatable sampling choice for the walkthrough; training can sample different windows.</p>\n<p><code>[32, 36, 40, 44, 48, 52, 56, 60, 64, 68, 72, 76, 80, 84, 88, 92]</code>\nFrame numbers start at zero. We&#39;ll read the selected frames in order and convert OpenCV&#39;s BGR colour order to RGB for display. Each item in <code>sampled_frames</code> will still be a complete 240 × 320 image, with three colour channels per pixel.</p>\n<p><code>(16, (240, 320, 3))</code>\nRead the grid from left to right, one row at a time. The players move towards the basket as the sequence progresses. These are the source frames, including the black borders and broadcast graphics already present in the clip.</p>\n<p>The pair labels connect this sequence to the model&#39;s input: frames 32 and 36 form the first pair, frames 40 and 44 the second, and so on. After spatial preprocessing to 224 × 224, each pair supplies 196 tubelets. Each tubelet covers one fixed 16 × 16 region across the two frames, so a whole frame pair produces 196 tokens, not one.</p>\n<p>We now have the 16 RGB frames. Next we&#39;ll prepare their pixel values and dimensions for the pretrained encoder.</p>\n<h2 id=\"getting-the-clip-ready-for-the-encoder\">Getting the Clip Ready for the Encoder</h2>\n<p>Our frames are 240 × 320 pixels. We&#39;ll resize the shorter side to 256, preserving the aspect ratio, then take the central 224 × 224 region. Every frame gets the same crop, so preprocessing does not introduce artificial camera movement.</p>\n<p>This follows the single-view evaluation recipe in <a href=\"https://github.com/facebookresearch/jepa/blob/main/evals/video_classification_frozen/utils.py\" rel=\"nofollow ugc noopener\">Meta&#39;s implementation</a>. We&#39;ll write the steps out using OpenCV so we can inspect them; small interpolation differences mean this is not a bit-for-bit reproduction of that implementation.</p>\n<p><code>(16, 224, 224, 3)</code>\nThe crop removes some of the scene, so it is worth checking what remains before interpreting any model output. The array now has shape <code>(16, 224, 224, 3)</code>: frames, height, width, RGB channels.</p>\n<p>We&#39;ll divide the pixel values by 255, then subtract a fixed mean and divide by a fixed standard deviation for each colour channel. The constants below come from the same evaluation recipe. They are not statistics calculated from this basketball clip.</p>\n<p>Finally, we&#39;ll put the axes in the order the encoder expects: <code>(batch, channels, time, height, width)</code>. A batch is a group of clips processed together; ours contains just one.</p>\n<p><code>permute</code> rearranges the axes, preserving the frame order. <code>unsqueeze(0)</code> adds the batch axis. The result is a PyTorch tensor, a multidimensional array ready for the model. These values still represent pixels. The encoder&#39;s patch-embedding layer will turn them into tubelet embeddings.</p>\n<p><code>(torch.Size([1, 3, 16, 224, 224]), torch.float32)</code></p>\n<h2 id=\"a-first-look-at-the-target-encoder-s-learned-features\">A First Look at the Target Encoder&#39;s Learned Features</h2>\n<p>We can now pass the complete clip through V-JEPA&#39;s pretrained <strong>target encoder</strong>. The target encoder should return a 1,024-value vector at each of the clip&#39;s 1,568 tubelet positions. Attention lets each position gather information from across the clip, so these vectors describe more than isolated pixels. They are features, not action labels; recognising &quot;basketball dunk&quot; would still require a suitable classifier.</p>\n<p>We&#39;ll run this target-encoder-only feature extraction on a Modal GPU and return the results. First we need the Modal Python package and an authenticated connection.</p>\n<p><code>Modal v1.5.5: connection ready</code>\nA checkpoint stores learned weights. We&#39;ll use <code>vitl16.pth.tar</code> from <a href=\"https://github.com/facebookresearch/jepa#model-zoo\" rel=\"nofollow ugc noopener\">Meta&#39;s model zoo</a>. The file contains weights for the context encoder (<code>encoder</code>), target encoder (<code>target_encoder</code>) and predictor, but this walkthrough selects <strong>only</strong> <code>checkpoint[&quot;target_encoder&quot;]</code>, following the default in the <a href=\"https://github.com/facebookresearch/jepa/blob/main/evals/video_classification_frozen/eval.py\" rel=\"nofollow ugc noopener\">evaluation code</a>.</p>\n<p>During V-JEPA pretraining, the target encoder is the moving-average encoder that supplies representation targets. Here, training is over: we use that target encoder by itself as a frozen feature extractor. The context encoder and predictor remain in the checkpoint but are neither instantiated nor run in this section.</p>\n<p>The following code describes the remote software environment. Modal caches this environment, including the downloaded checkpoint, so subsequent runs can reuse it.</p>\n<p>Before running the clip, we&#39;ll require every <strong>target-encoder</strong> weight to match the <code>vit_large</code> model definition. <code>strict=True</code> makes loading fail if any selected target-encoder weights are missing or unexpected, instead of silently leaving part of that encoder randomly initialised.</p>\n<p>We&#39;ll also return the checkpoint&#39;s top-level keys to verify which components the file contains for later experiments. Seeing <code>encoder</code>, <code>target_encoder</code> and <code>predictor</code> in that list does not mean all three are used here: the loader extracts only <code>checkpoint[&quot;target_encoder&quot;]</code>.</p>\n<p><code>eval()</code> selects evaluation behaviour, while <code>inference_mode()</code> avoids building the gradient information needed for training. We pass the complete clip directly to the frozen target encoder, without a mask. It produces one feature vector at every tubelet position, and no weights are updated.</p>\n<p>Nothing in this call uses V-JEPA&#39;s context encoder or predictor. Those components are central to the pretraining task, where visible context is encoded and missing-region representations are predicted, but this is inference after pretraining, using only the learned target encoder to describe the clip.</p>\n<p>The first call also builds the remote environment and downloads the checkpoint, so it will take longer than a cached run. This uses one A10G GPU on Modal for the inference call.</p>\n<pre><code>Checkpoint contains: encoder, target_encoder, predictor\nThis run used: target_encoder\nEncoded clip: (1, 1568, 1024)\nJEPA revision: 51c59d51</code></pre>\n<p>The returned shape is <code>(1, 1568, 1024)</code>: one clip, 1,568 target-encoder positions and 1,024 values per position. Strict loading passed for the selected target-encoder weights.</p>\n<p>The checkpoint also contains <code>encoder</code> and <code>predictor</code> weights, but their presence should not be confused with execution: this run loaded and called only <code>target_encoder</code>. No context-encoder or predictor outputs contribute to <code>features</code>.</p>\n<p>We can arrange the 1,568 positions back into the familiar grid of eight frame pairs, 14 rows and 14 columns. That only changes how we index the target encoder&#39;s output. Each vector already includes information gathered from across the clip through the target encoder&#39;s attention layers.</p>\n<pre><code>((8, 14, 14, 1024),\n array([ 0.025, -0.954, -0.669,  1.332,  0.463, -0.93 , -0.63 ,  0.265,\n         0.721, -0.162, -0.434,  0.407], dtype=float32))</code></pre>\n<p>Those are the first 12 values of one 1,024-value <strong>target-encoder</strong> vector. Row and column indices start at zero. Individual coordinates do not come with labels such as &quot;ball&quot; or &quot;player&quot;, so this list alone tells us little about what the target encoder has learned. Comparing its vectors across positions or under controlled changes to the clip will be more informative.</p>\n<p>These saved features are target-encoder representations only; they contain no separately computed context-encoder representation and no predictor output. We can save the complete features and the source-code revision here, so inspecting this result again does not require another GPU run.</p>\n<h2 id=\"what-changes-when-time-runs-differently\">What Changes When Time Runs Differently?</h2>\n<p>The previous experiment used spatial masking while leaving the clip&#39;s timeline intact. We can now hold the 16 sampled images fixed and change only their order. If V-JEPA represented a clip as an unordered collection of pictures, these changes would make no difference. Its tubelets and temporal position vectors give us reason to expect otherwise.</p>\n<p>We&#39;ll make two altered clips. One swaps the two frames inside every tubelet pair. The other reverses the complete sequence. Both contain exactly the same normalised frames as the original, with no new pixels and none removed.</p>\n<p><code>{&#39;original&#39;: [32, 36, 40, 44, 48, 52, 56, 60, 64, 68, 72, 76, 80, 84, 88, 92], &#39;reversed&#39;: [92, 88, 84, 80, 76, 72, 68, 64, 60, 56, 52, 48, 44, 40, 36, 32], &#39;within-pair swap&#39;: [36, 32, 44, 40, 52, 48, 60, 56, 68, 64, 76, 72, 84, 80, 92, 88]}</code>\nThe within-pair version keeps each pair in the same temporal slot but reverses its two frames: <code>(32, 36)</code> becomes <code>(36, 32)</code>. This tests whether a two-frame tubelet responds to local direction.</p>\n<p>The fully reversed version also moves the pairs to different temporal slots. When we compare its outputs with the original, we&#39;ll reverse the eight output groups back into source-pair order. Matching positions will then refer to the same two source frames and spatial patch, although the model encountered them in the opposite temporal direction.</p>\n<p><code>(torch.Size([1, 3, 16, 224, 224]), torch.Size([1, 3, 16, 224, 224]))</code>\nWe&#39;ll use the same frozen target encoder and the same preprocessing as before. Each clip is encoded separately, so the GPU does not mix information between variants. The output for each one remains <code>(1, 1568, 1024)</code>.</p>\n<p><code>(3, 1568, 1024)</code>\nCosine similarity compares the direction of two feature vectors while ignoring their overall length. A value of 1 means identical direction, 0 means no directional alignment, and negative values point in opposing directions.</p>\n<p>We&#39;ll calculate it for matching spatial positions and source-frame pairs, then average over the 196 positions in each pair. For the fully reversed clip, this requires reversing the eight output groups before comparison.</p>\n<p><code>(0.6825879216194153, 0.4968155026435852)</code>\nSwapping the two frames inside every tubelet pair gives a mean cosine similarity of about 0.683. The source images, pair membership and temporal slots are unchanged, so the drop from 1 shows that frame order inside a tubelet affects the encoder&#39;s features.</p>\n<p>Reversing the complete clip lowers the aligned mean similarity to about 0.497. That change includes the reversal inside each pair, different temporal positions for the pairs, and attention across the reversed sequence. This experiment shows sensitivity to temporal order on one clip. It does not establish that the encoder recognises the action, understands cause and effect, or would respond the same way across a dataset.</p>\n<p>The stronger response to complete reversal is consistent with a representation that uses information over time as well as appearance. It is not possible to assign that extra change to one mechanism from this comparison alone, because temporal position and cross-token attention change together.</p>\n<p>The next section will ask a more practical question: can a small classifier use frozen V-JEPA features to separate actions without changing the encoder?</p>\n<p><code>Saved vjepa_temporal_order.npz</code>\nAs with the earlier feature-extraction and hidden-region experiments, we save these results locally in SolveIt as a compressed NumPy archive. This avoids repeating the GPU inference when we return to the temporal-order comparison.</p>\n<p>This archive records the target-encoder features for the original, within-pair-swapped and fully reversed clips; the corresponding source-frame orders; the position-wise cosine similarities; and the JEPA source revision used for the run. Keeping the orders and revision beside the arrays makes the comparison easier to interpret and reproduce.</p>\n<p>It can be reopened with <code>np.load(&quot;vjepa_temporal_order.npz&quot;)</code>. Its contents are then available by name, including <code>data[&quot;features&quot;]</code>, <code>data[&quot;pair_swap_cosine&quot;]</code> and <code>data[&quot;reverse_cosine&quot;]</code>.</p>\n<h2 id=\"can-frozen-v-jepa-features-separate-actions\">Can Frozen V-JEPA Features Separate Actions?</h2>\n<p>Pretraining tells us how V-JEPA learns, but the practical question is what its encoder is useful for afterwards. We&#39;ll test that with a <strong>linear probe</strong>: freeze the released target encoder, reduce each clip to one feature vector, and train only a linear classifier to distinguish three actions.</p>\n<p>The experiment has four stages:</p>\n<ol><li>choose clips from <code>Archery</code> ,<code>BabyCrawling</code> , and<code>BasketballDunk</code> ;</li><li>preprocess each clip into the same 16-frame input used above;</li><li>average the encoder&#39;s 1,568 output tokens into one 1,024-value clip representation; and</li><li>fit a three-way linear classifier on those frozen representations.</li></ol>\n<p>A linear probe is deliberately limited. If it separates the actions, the class information was already accessible in the pretrained features; the classifier did not teach the encoder a new representation. We use the target encoder here for consistency with the preceding feature experiments and Meta&#39;s frozen-evaluation default.</p>\n<p>This is a small demonstration rather than a UCF101 benchmark: only 18 clips are used, from three visually distinct classes. To make the test more meaningful, we split by UCF101 <strong>recording group</strong> rather than by individual clip. Four groups per class are used for training and two different groups for testing, so near-duplicate clips from the same original recording cannot appear on both sides.</p>\n<p>The probe reuses the Modal app, checkpoint image, and strict target-encoder loader from the earlier experiments. The only additional remote dependency is OpenCV, which decodes the UCF101 videos. The next cell adds OpenCV to that environment; the following check spells out the expected clip shape and resulting tubelet count.</p>\n<p><code>((3, 16, 224, 224), (2, 16, 16), 1568)</code></p>\n<h3 id=\"selecting-clips-and-splitting-by-recording-group\">Selecting Clips and Splitting by Recording Group</h3>\n<p>The downloaded UCF101 subset contains several clips from each recording group. We keep one clip per group, then assign the first four groups in each class to training and the next two to testing. The important unit of separation is the group, not the filename: clips cut from one recording stay on only one side of the split.</p>\n<p>Before preprocessing the clips, we can inspect the exact sample used by the probe. Each row below is one selected recording-group clip; the five columns show evenly spaced source frames from beginning to end. The row labels identify the action and whether the clip belongs to the training or held-out split.</p>\n<p>This is a useful visual check for near-duplicate scenes, uninformative opening or closing frames, and background cues that might make the small classification task easier than the action itself.</p>\n<h3 id=\"turning-each-video-into-one-model-input\">Turning Each Video into One Model Input</h3>\n<p>The source clips vary in length, so we choose 16 evenly spaced frames from the beginning to the end of each clip. Each frame then follows the same resize, centre-crop and channel-normalisation recipe used earlier.</p>\n<p>This sampling choice is intentionally simple and deterministic. It gives every video the required shape, but it is not the multi-view evaluation protocol used for a full benchmark.</p>\n<h3 id=\"extracting-frozen-clip-representations\">Extracting Frozen Clip Representations</h3>\n<p>For each selected video, the target encoder produces 1,568 vectors. We average them across space and time to obtain one 1,024-value vector for the complete clip. This pooling discards the location of individual tubelets, but leaves a compact representation suitable for a small classifier.</p>\n<p>The encoder stays in evaluation mode and no gradients are computed. The action labels are collected alongside the features, but they are never supplied to V-JEPA.</p>\n<p>With the selection and preprocessing functions in place, we can run feature extraction once on the GPU. The returned arrays contain the pooled features, labels, split assignments and filenames needed for the local probe; the classifier itself will run in SolveIt rather than on the GPU.</p>\n<pre><code>((18, 1024),\n np.int64(12),\n np.int64(6),\n [&#39;Archery&#39;, &#39;BabyCrawling&#39;, &#39;BasketballDunk&#39;])</code></pre>\n<h3 id=\"training-the-linear-probe\">Training the Linear Probe</h3>\n<p>The extractor returns 18 frozen clip vectors, each 1,024 values wide. For every class, four UCF101 recording groups supply the training examples and two different groups supply the test examples. That gives us 12 training clips and 6 held-out clips, with no recording group shared across the split.</p>\n<p>The classifier is deliberately small. We normalise every clip vector to unit length, append a constant bias feature, and fit a regularised linear map from the 1,024 encoder features to three class scores. The highest score becomes the predicted action.</p>\n<p>There are only 12 training examples but more than 1,000 feature coordinates, so regularisation matters: without it, many solutions could memorise the training set. The expression below is ridge regression written in its dual form, which solves a 12 × 12 system instead of a 1,025 × 1,025 one. Only these classifier weights learn from the labels, while V-JEPA remains frozen.</p>\n<pre><code>(np.float64(1.0),\n np.float64(1.0),\n [(np.int64(0), np.int64(0)),\n  (np.int64(0), np.int64(0)),\n  (np.int64(1), np.int64(1)),\n  (np.int64(1), np.int64(1)),\n  (np.int64(2), np.int64(2)),\n  (np.int64(2), np.int64(2))])</code></pre>\n<p>The feature matrix first converts to 64-bit floating point for a stable linear solve. Each 1,024-value clip vector is divided by its length, so the classifier compares the direction of the representation rather than allowing vectors with larger magnitudes to dominate. Appending a constant value of one gives the linear classifier a bias term.</p>\n<p>The training labels are converted to one-hot target vectors: for example, an Archery clip has target <code>[1, 0, 0]</code>. Ridge regression then learns a regularised linear mapping from the frozen V-JEPA representations to three class scores. The regularisation strength is <code>0.1</code>; it discourages excessively large classifier weights in this setting with many more feature coordinates than training examples.</p>\n<p>For each clip, <code>argmax</code> selects the class with the highest score. Both reported accuracies are <code>1.0</code>, meaning that the classifier correctly labels all 12 training clips and all 6 held-out clips. The <code>(expected, predicted)</code> pairs confirm that every held-out numeric label matches its prediction: class <code>0</code> is Archery, <code>1</code> is BabyCrawling and <code>2</code> is BasketballDunk.</p>\n<p>Perfect training accuracy is not surprising with so few examples. The held-out result is more informative because its clips come from different recording groups, but six test clips are far too few to estimate performance on UCF101 generally.</p>\n<pre><code>v_Archery_g05_c04.avi                  Archery        -&gt; Archery\nv_Archery_g06_c01.avi                  Archery        -&gt; Archery\nv_BabyCrawling_g05_c01.avi             BabyCrawling   -&gt; BabyCrawling\nv_BabyCrawling_g06_c04.avi             BabyCrawling   -&gt; BabyCrawling\nv_BasketballDunk_g05_c01.avi           BasketballDunk -&gt; BasketballDunk\nv_BasketballDunk_g06_c01.avi           BasketballDunk -&gt; BasketballDunk\nHeld-out accuracy: 100% (chance: 33%)</code></pre>\n<p>The printed rows translate the numeric predictions back into class names. Each row shows the held-out filename, its expected action and the classifier&#39;s prediction. All six predictions are correct: two Archery clips, two BabyCrawling clips and two BasketballDunk clips.</p>\n<p>The resulting held-out accuracy is therefore $6/6 = 100\\%$, compared with a 33% chance level for three balanced classes. This demonstrates that these three actions are linearly separable in this small sample of frozen V-JEPA features. It does not establish 100% accuracy on new videos or on the full 101-class dataset; the actions are visually distinct and the test set is intentionally small.</p>\n<p>Finally, <code>vjepa_linear_probe.npz</code> saves the frozen features, labels, split assignments, filenames, learned classifier weights and class names. This lets us inspect or reuse the probe without downloading the videos and running the encoder again.</p>\n<h3 id=\"reading-the-held-out-result\">Reading the Held-Out Result</h3>\n<p>All six held-out predictions are correct. The confusion matrix below shows their distribution across the three classes.</p>\n<p>Rows are the true classes, columns are the predicted classes, and each cell counts clips. A perfect result places all six clips on the diagonal.</p>\n<p>The linear probe correctly classifies all six held-out clips, compared with a random 33% chance level for three classes. Its input is only the frozen 1,024-value clip vector. The V-JEPA encoder is never updated, so the result gives us direct evidence that its representation already separates these actions in this small sample.</p>\n<p>The sample is far too small for a performance claim. Archery, baby crawling and basketball dunking are visually distinct, and six test clips cannot represent the variation in UCF101. <strong>The useful point is narrower: a shallow supervised head can read action information from features learned without action labels.</strong></p>\n<p>The confusion matrix is completely diagonal, matching the six printed predictions. More data and more classes would be needed to test how robust this separation is, but the experiment answers our immediate question: a simple linear classifier can distinguish these three actions from frozen V-JEPA features in this small held-out sample. The encoder itself was not retrained.</p>\n<h2 id=\"summary\">Summary</h2>\n<p>We began with raw video: 16 sampled frames from a basketball clip. V-JEPA divided them into 1,568 space-time tubelets and represented each one with a vector of 1,024 learned features. Unlike a pixel-reconstruction model, V-JEPA was trained to predict these representations for hidden parts of a video from the parts it could see.</p>\n<p>Our experiments examined that idea from three angles. First, the pretrained predictor estimated the target encoder&#39;s features for a withheld central region more accurately than a zero-vector baseline. Changing the hidden pixels had no effect on those predictions, confirming that their contents had not leaked into the visible context.</p>\n<p>Next, we changed the order of the same 16 frames. Swapping frames within each tubelet pair changed the encoder&#39;s features, and reversing the complete clip changed them more strongly. On this example, V-JEPA&#39;s representation therefore depended on temporal order rather than treating the video as an unordered collection of images.</p>\n<p>Finally, we froze the encoder and trained only a simple linear classifier on its clip-level features. It correctly separated archery, baby crawling and basketball dunk in a small test set drawn from held-out recording groups. This does not constitute a UCF101 benchmark, but it shows that useful action information was already accessible in the pretrained representation.</p>\n<p>Together, these results illustrate the main idea behind V-JEPA: learning useful video representations by predicting in feature space rather than reconstructing pixels. The experiments are deliberately small: one mask, one temporal-order example and three action classes, but they make the model&#39;s data flow and capabilities concrete.</p>\n<p>For the formal method and full evaluations, see the <a href=\"https://arxiv.org/abs/2404.08471\" rel=\"nofollow ugc noopener\">V-JEPA paper</a> and the <a href=\"https://github.com/facebookresearch/jepa\" rel=\"nofollow ugc noopener\">official JEPA repository</a>.</p>\n<p>Community</p>\n<h2 id=\"explore-together\">Explore together.</h2>\n<p>Share what you&#39;re learning, ask questions, and swap ideas.</p>\n<p> Join the ExploringML Discord.</p>\n<p><a href=\"https://discord.gg/wXhHXc9puA\" rel=\"nofollow ugc noopener\">Join the community</a></p>","headings":[{"level":1,"text":"V-JEPA: Learning Video Representations by Feature Prediction","id":"v-jepa-learning-video-representations-by-feature-prediction"},{"level":2,"text":"One Pretraining Step","id":"one-pretraining-step"},{"level":2,"text":"What Changes When the Picture Moves?","id":"what-changes-when-the-picture-moves"},{"level":2,"text":"The Model We'll Explore","id":"the-model-we-ll-explore"},{"level":3,"text":"Count the Video Patches","id":"count-the-video-patches"},{"level":2,"text":"Choosing the Frames","id":"choosing-the-frames"},{"level":2,"text":"Getting the Clip Ready for the Encoder","id":"getting-the-clip-ready-for-the-encoder"},{"level":2,"text":"A First Look at the Target Encoder's Learned Features","id":"a-first-look-at-the-target-encoder-s-learned-features"},{"level":2,"text":"What Changes When Time Runs Differently?","id":"what-changes-when-time-runs-differently"},{"level":2,"text":"Can Frozen V-JEPA Features Separate Actions?","id":"can-frozen-v-jepa-features-separate-actions"},{"level":3,"text":"Selecting Clips and Splitting by Recording Group","id":"selecting-clips-and-splitting-by-recording-group"},{"level":3,"text":"Turning Each Video into One Model Input","id":"turning-each-video-into-one-model-input"},{"level":3,"text":"Extracting Frozen Clip Representations","id":"extracting-frozen-clip-representations"},{"level":3,"text":"Training the Linear Probe","id":"training-the-linear-probe"},{"level":3,"text":"Reading the Held-Out Result","id":"reading-the-held-out-result"},{"level":2,"text":"Summary","id":"summary"},{"level":2,"text":"Explore together.","id":"explore-together"}]}}