A HuggingFace-format conversion of Meta AI's
V-JEPA 2.1 ViT-B/16
video encoder and predictor,
operating at 384x384 resolution.
The weights are Meta's, copied without modification. This repository provides the
transformers
-compatible packaging plus a documented numerical validation against the
original implementation.
An equivalent community port already exists (
Dev-Jahn/vjepa2.1-vitb-fpc64-384
). This repository adds an independently reproduced conversion together with the validation results below.
modeling_vjepa21.py
/
configuration_vjepa21.py
, same repository
Weight modifications
None. Every tensor is a bit-exact copy of the original.
The only structural change is that the fused QKV projection of each attention block is split into
separate
query
/
key
/
value
matrices, following the convention used by
transformers
.
This is a re-parameterization, not a change of weights.
Architecture
Encoder
Predictor
Hidden size
768
384
Layers
12
12
Attention heads
12
12
MLP ratio
4.0
4.0
Patch size
16 x 16, tubelet 2
—
Input resolution
384 x 384
—
Position encoding
3D RoPE
3D RoPE
Parameters
86.8 M
22.9 M
Total:
109.7 M
.
Distilled from ViT-G. The predictor projects to a 1664-dim teacher space.
The encoder carries a separate patch embedding for single images, learnable image/video modality
embeddings, and four per-layer normalizations used for hierarchical feature extraction.
Usage
import torch
from transformers import AutoModel
model = AutoModel.from_pretrained(
"apiantonio/vjepa2.1-vitb-384",
trust_remote_code=True,
dtype=torch.float32,
).eval().cuda()
# (batch, channels, frames, height, width) — channels-first, following the original# implementation. Note this differs from the (batch, frames, channels, H, W) layout# used by the transformers V-JEPA 2 (2.0) classes.
video = torch.randn(1, 3, 64, 384, 384).cuda()
with torch.no_grad():
out = model(video, skip_predictor=True)
feats = out.last_hidden_state # (1, 18432, 768) for 64 frames
Token grid:
(frames / 2) x 24 x 24
. A 64-frame clip yields 32 x 24 x 24 = 18,432 tokens.
Frames are expected to be normalized. See "Not verified" below before relying on a specific
normalization; the safest reference is the transform pipeline in the official repository.
Validation
Two independent checks were run against the original implementation. All tests use float32
and
torch.no_grad()
.
1. Weight-level provenance
Every tensor in the published
safetensors
was traced back to a tensor in the official
checkpoint and compared element-wise.
Check
Result
max|Δ|
across all tensors
0.000e+00
Parameters with no origin in the checkpoint
0
The second row matters: the conversion script calls
load_state_dict(..., strict=False)
, which
silently leaves unmatched parameters randomly initialized. Zero orphans means no such parameter
survived — for the predictor as well as the encoder.
2. Forward parity
The same input tensor was passed through Meta's official encoder (loaded via
torch.hub
) and
through this port, at three clip lengths.
T
Tokens
max|Δ|
mean|Δ|
mean|Δ| / mean|a|
cos-sim (min)
16
4,608
0.000e+00
0.000e+00
0.000e+00
1.000000
32
9,216
0.000e+00
0.000e+00
0.000e+00
0.999999*
64
18,432
0.000e+00
0.000e+00
0.000e+00
1.000000
*
cosine_similarity
computes
a·b / (‖a‖·‖b‖)
; numerator and denominator follow different rounding paths, so the metric can deviate in the last digit even for bit-identical tensors. The absolute differences are exactly zero.
Summary.
Bit-exact agreement with the reference implementation at every tested clip length. Varying T from 16 to 64 exercises the RoPE position interpolation,
which is the most likely silent failure mode in a port of this architecture; no growth of the
residual with sequence length was observed.
Not verified
Stated explicitly so the scope of the validation above is not overread:
Predictor forward pass.
All parity tests use
skip_predictor=True
. The predictor
weights
are verified bit-exact, but its forward logic (mask-token injection, context projection) has not
been compared against the reference.
Video preprocessing.
Normalization constants and frame sampling were not validated against
the official transform pipeline. No
VideoProcessor
is shipped with this repository.
Downstream benchmarks.
No evaluation numbers from the V-JEPA 2.1 release were reproduced.
This port is validated for numerical equivalence, not for task performance.
License and attribution
The weights are released by Meta AI under the license of the original V-JEPA 2 release (MIT).
This repository redistributes them unmodified and adds no additional restrictions.
Credit where it is due:
Meta AI
— the V-JEPA 2 / V-JEPA 2.1 research and the pretrained weights.
Dev-Jahn
— the
transformers
-compatible modeling
code and the conversion script. Both are used here essentially unchanged.
The contribution specific to this repository is the conversion of the B variant and the
numerical validation documented above.
Citation
@article{assran2025vjepa2,
title={V-JEPA~2: Self-Supervised Video Models Enable Understanding, Prediction and Planning},
author={Assran, Mahmoud and Bardes, Adrien and Fan, David and Garrido, Quentin and Howes, Russell and
Komeili, Mojtaba and Muckley, Matthew and Rizvi, Ammar and Roberts, Claire and Sinha, Koustuv and Zholus, Artem and
Arnaud, Sergio and Gejji, Abha and Martin, Ada and Robert Hogan, Francois and Dugas, Daniel and
Bojanowski, Piotr and Khalidov, Vasil and Labatut, Patrick and Massa, Francisco and Szafraniec, Marc and
Krishnakumar, Kapil and Li, Yong and Ma, Xiaodong and Chandar, Sarath and Meier, Franziska and LeCun, Yann and
Rabbat, Michael and Ballas, Nicolas},
journal={arXiv preprint arXiv:2506.09985},
year={2025}
}
@article{murlabadia2026vjepa2_1,
title={V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning},
author={Mur-Labadia, Lorenzo and Muckley, Matthew and Bar, Amir and Assran, Mahmoud and
Sinha, Koustuv and Rabbat, Michael and LeCun, Yann and Ballas, Nicolas and Bardes, Adrien},
journal={arXiv preprint arXiv:2603.14482},
year={2026}
}
Please cite the original V-JEPA 2 work. If the format conversion itself was useful, a link back to
Dev-Jahn/vjepa2-hf
is appreciated.
Runs of apiantonio vjepa2.1-vit-base-384 on huggingface.co
3.4K
Total runs
-20
24-hour runs
-415
3-day runs
-708
7-day runs
-971
30-day runs
More Information About vjepa2.1-vit-base-384 huggingface.co Model
vjepa2.1-vit-base-384 huggingface.co is an AI model on huggingface.co that provides vjepa2.1-vit-base-384's model effect (), which can be used instantly with this apiantonio vjepa2.1-vit-base-384 model. huggingface.co supports a free trial of the vjepa2.1-vit-base-384 model, and also provides paid use of the vjepa2.1-vit-base-384. Support call vjepa2.1-vit-base-384 model through api, including Node.js, Python, http.
vjepa2.1-vit-base-384 huggingface.co is an online trial and call api platform, which integrates vjepa2.1-vit-base-384's modeling effects, including api services, and provides a free online trial of vjepa2.1-vit-base-384, you can try vjepa2.1-vit-base-384 online for free by clicking the link below.
apiantonio vjepa2.1-vit-base-384 online free url in huggingface.co:
vjepa2.1-vit-base-384 is an open source model from GitHub that offers a free installation service, and any user can find vjepa2.1-vit-base-384 on GitHub to install. At the same time, huggingface.co provides the effect of vjepa2.1-vit-base-384 install, users can directly use vjepa2.1-vit-base-384 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
vjepa2.1-vit-base-384 install url in huggingface.co: