A frontier video understanding model developed by FAIR, Meta, which extends the pretraining objectives of
VJEPA
, resulting in state-of-the-art video understanding capabilities, leveraging data and model sizes at scale.
The code is released
in this repository
.
Installation
To run V-JEPA 2 model, ensure you have installed the latest transformers:
V-JEPA 2 is intended to represent any video (and image) to perform video classification, retrieval, or as a video encoder for VLMs.
from transformers import AutoVideoProcessor, AutoModel
hf_repo = "facebook/vjepa2-vitg-fpc64-256"
model = AutoModel.from_pretrained(hf_repo)
processor = AutoVideoProcessor.from_pretrained(hf_repo)
To load a video, sample the number of frames according to the model. For this model, we use 64.
import torch
from torchcodec.decoders import VideoDecoder
import numpy as np
video_url = "https://huggingface.co/datasets/nateraw/kinetics-mini/resolve/main/val/archery/-Qz25rXdMjE_000014_000024.mp4"
vr = VideoDecoder(video_url)
frame_idx = np.arange(0, 64) # choosing some frames. here, you can define more complex sampling strategy
video = vr.get_frames_at(indices=frame_idx).data # T x C x H x W
video = processor(video, return_tensors="pt").to(model.device)
with torch.no_grad():
video_embeddings = model.get_vision_features(**video)
print(video_embeddings.shape)
To load an image, simply copy the image to the desired number of frames.
For more code examples, please refer to the V-JEPA 2 documentation.
Citation
@techreport{assran2025vjepa2,
title={V-JEPA~2: Self-Supervised Video Models Enable Understanding, Prediction and Planning},
author={Assran, Mahmoud and Bardes, Adrien and Fan, David and Garrido, Quentin and Howes, Russell and
Komeili, Mojtaba and Muckley, Matthew and Rizvi, Ammar and Roberts, Claire and Sinha, Koustuv and Zholus, Artem and
Arnaud, Sergio and Gejji, Abha and Martin, Ada and Robert Hogan, Francois and Dugas, Daniel and
Bojanowski, Piotr and Khalidov, Vasil and Labatut, Patrick and Massa, Francisco and Szafraniec, Marc and
Krishnakumar, Kapil and Li, Yong and Ma, Xiaodong and Chandar, Sarath and Meier, Franziska and LeCun, Yann and
Rabbat, Michael and Ballas, Nicolas},
institution={FAIR at Meta},
year={2025}
}
Runs of facebook vjepa2-vitg-fpc64-256 on huggingface.co
88.7K
Total runs
0
24-hour runs
-4.3K
3-day runs
-17.2K
7-day runs
-52.5K
30-day runs
More Information About vjepa2-vitg-fpc64-256 huggingface.co Model
vjepa2-vitg-fpc64-256 huggingface.co is an AI model on huggingface.co that provides vjepa2-vitg-fpc64-256's model effect (), which can be used instantly with this facebook vjepa2-vitg-fpc64-256 model. huggingface.co supports a free trial of the vjepa2-vitg-fpc64-256 model, and also provides paid use of the vjepa2-vitg-fpc64-256. Support call vjepa2-vitg-fpc64-256 model through api, including Node.js, Python, http.
vjepa2-vitg-fpc64-256 huggingface.co is an online trial and call api platform, which integrates vjepa2-vitg-fpc64-256's modeling effects, including api services, and provides a free online trial of vjepa2-vitg-fpc64-256, you can try vjepa2-vitg-fpc64-256 online for free by clicking the link below.
facebook vjepa2-vitg-fpc64-256 online free url in huggingface.co:
vjepa2-vitg-fpc64-256 is an open source model from GitHub that offers a free installation service, and any user can find vjepa2-vitg-fpc64-256 on GitHub to install. At the same time, huggingface.co provides the effect of vjepa2-vitg-fpc64-256 install, users can directly use vjepa2-vitg-fpc64-256 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
vjepa2-vitg-fpc64-256 install url in huggingface.co: