apiantonio / vjepa2.1-vit-base-384

huggingface.co
Total runs: 3.4K
24-hour runs: -20
7-day runs: -708
30-day runs: -971
Model's Last Updated: July 27 2026
video-classification

Introduction of vjepa2.1-vit-base-384

Model Details of vjepa2.1-vit-base-384

V-JEPA 2.1 ViT-B/16 384 — HuggingFace port

A HuggingFace-format conversion of Meta AI's V-JEPA 2.1 ViT-B/16 video encoder and predictor, operating at 384x384 resolution.

The weights are Meta's, copied without modification. This repository provides the transformers -compatible packaging plus a documented numerical validation against the original implementation.

An equivalent community port already exists ( Dev-Jahn/vjepa2.1-vitb-fpc64-384 ). This repository adds an independently reproduced conversion together with the validation results below.

Provenance
Original weights Meta AI — facebookresearch/vjepa2
Source checkpoint vjepa2_1_vitb_dist_vitG_384.pt
State dict key ema_encoder
Conversion script convert_vjepa21_to_hf.py from Dev-Jahn/vjepa2-hf
Modeling code modeling_vjepa21.py / configuration_vjepa21.py , same repository
Weight modifications None. Every tensor is a bit-exact copy of the original.

The only structural change is that the fused QKV projection of each attention block is split into separate query / key / value matrices, following the convention used by transformers . This is a re-parameterization, not a change of weights.

Architecture
Encoder Predictor
Hidden size 768 384
Layers 12 12
Attention heads 12 12
MLP ratio 4.0 4.0
Patch size 16 x 16, tubelet 2 —
Input resolution 384 x 384 —
Position encoding 3D RoPE 3D RoPE
Parameters 86.8 M 22.9 M

Total: 109.7 M .

Distilled from ViT-G. The predictor projects to a 1664-dim teacher space.

The encoder carries a separate patch embedding for single images, learnable image/video modality embeddings, and four per-layer normalizations used for hierarchical feature extraction.

Usage
import torch
from transformers import AutoModel

model = AutoModel.from_pretrained(
    "apiantonio/vjepa2.1-vitb-384",
    trust_remote_code=True,
    dtype=torch.float32,
).eval().cuda()

# (batch, channels, frames, height, width) — channels-first, following the original
# implementation. Note this differs from the (batch, frames, channels, H, W) layout
# used by the transformers V-JEPA 2 (2.0) classes.
video = torch.randn(1, 3, 64, 384, 384).cuda()

with torch.no_grad():
    out = model(video, skip_predictor=True)

feats = out.last_hidden_state       # (1, 18432, 768) for 64 frames

Token grid: (frames / 2) x 24 x 24 . A 64-frame clip yields 32 x 24 x 24 = 18,432 tokens.

Frames are expected to be normalized. See "Not verified" below before relying on a specific normalization; the safest reference is the transform pipeline in the official repository.

Validation

Two independent checks were run against the original implementation. All tests use float32 and torch.no_grad() .

1. Weight-level provenance

Every tensor in the published safetensors was traced back to a tensor in the official checkpoint and compared element-wise.

Check Result
max|Δ| across all tensors 0.000e+00
Parameters with no origin in the checkpoint 0

The second row matters: the conversion script calls load_state_dict(..., strict=False) , which silently leaves unmatched parameters randomly initialized. Zero orphans means no such parameter survived — for the predictor as well as the encoder.

2. Forward parity

The same input tensor was passed through Meta's official encoder (loaded via torch.hub ) and through this port, at three clip lengths.

T Tokens max|Δ| mean|Δ| mean|Δ| / mean|a| cos-sim (min)
16 4,608 0.000e+00 0.000e+00 0.000e+00 1.000000
32 9,216 0.000e+00 0.000e+00 0.000e+00 0.999999*
64 18,432 0.000e+00 0.000e+00 0.000e+00 1.000000

* cosine_similarity computes a·b / (‖a‖·‖b‖) ; numerator and denominator follow different rounding paths, so the metric can deviate in the last digit even for bit-identical tensors. The absolute differences are exactly zero.

Summary. Bit-exact agreement with the reference implementation at every tested clip length. Varying T from 16 to 64 exercises the RoPE position interpolation, which is the most likely silent failure mode in a port of this architecture; no growth of the residual with sequence length was observed.

Not verified

Stated explicitly so the scope of the validation above is not overread:

  • Predictor forward pass. All parity tests use skip_predictor=True . The predictor weights are verified bit-exact, but its forward logic (mask-token injection, context projection) has not been compared against the reference.
  • Video preprocessing. Normalization constants and frame sampling were not validated against the official transform pipeline. No VideoProcessor is shipped with this repository.
  • Downstream benchmarks. No evaluation numbers from the V-JEPA 2.1 release were reproduced. This port is validated for numerical equivalence, not for task performance.
License and attribution

The weights are released by Meta AI under the license of the original V-JEPA 2 release (MIT). This repository redistributes them unmodified and adds no additional restrictions.

Credit where it is due:

  • Meta AI — the V-JEPA 2 / V-JEPA 2.1 research and the pretrained weights.
  • Dev-Jahn — the transformers -compatible modeling code and the conversion script. Both are used here essentially unchanged.

The contribution specific to this repository is the conversion of the B variant and the numerical validation documented above.

Citation
@article{assran2025vjepa2,
  title={V-JEPA~2: Self-Supervised Video Models Enable Understanding, Prediction and Planning},
  author={Assran, Mahmoud and Bardes, Adrien and Fan, David and Garrido, Quentin and Howes, Russell and
Komeili, Mojtaba and Muckley, Matthew and Rizvi, Ammar and Roberts, Claire and Sinha, Koustuv and Zholus, Artem and
Arnaud, Sergio and Gejji, Abha and Martin, Ada and Robert Hogan, Francois and Dugas, Daniel and
Bojanowski, Piotr and Khalidov, Vasil and Labatut, Patrick and Massa, Francisco and Szafraniec, Marc and
Krishnakumar, Kapil and Li, Yong and Ma, Xiaodong and Chandar, Sarath and Meier, Franziska and LeCun, Yann and
Rabbat, Michael and Ballas, Nicolas},
  journal={arXiv preprint arXiv:2506.09985},
  year={2025}
}
@article{murlabadia2026vjepa2_1,
  title={V-JEPA 2.1: Unlocking Dense Features in Video Self-Supervised Learning},
  author={Mur-Labadia, Lorenzo and Muckley, Matthew and Bar, Amir and Assran, Mahmoud and
Sinha, Koustuv and Rabbat, Michael and LeCun, Yann and Ballas, Nicolas and Bardes, Adrien},
  journal={arXiv preprint arXiv:2603.14482},
  year={2026}
}

Please cite the original V-JEPA 2 work. If the format conversion itself was useful, a link back to Dev-Jahn/vjepa2-hf is appreciated.

Runs of apiantonio vjepa2.1-vit-base-384 on huggingface.co

3.4K
Total runs
-20
24-hour runs
-415
3-day runs
-708
7-day runs
-971
30-day runs

More Information About vjepa2.1-vit-base-384 huggingface.co Model

More vjepa2.1-vit-base-384 license Visit here:

https://choosealicense.com/licenses/mit

vjepa2.1-vit-base-384 huggingface.co

vjepa2.1-vit-base-384 huggingface.co is an AI model on huggingface.co that provides vjepa2.1-vit-base-384's model effect (), which can be used instantly with this apiantonio vjepa2.1-vit-base-384 model. huggingface.co supports a free trial of the vjepa2.1-vit-base-384 model, and also provides paid use of the vjepa2.1-vit-base-384. Support call vjepa2.1-vit-base-384 model through api, including Node.js, Python, http.

vjepa2.1-vit-base-384 huggingface.co Url

https://huggingface.co/apiantonio/vjepa2.1-vit-base-384

apiantonio vjepa2.1-vit-base-384 online free

vjepa2.1-vit-base-384 huggingface.co is an online trial and call api platform, which integrates vjepa2.1-vit-base-384's modeling effects, including api services, and provides a free online trial of vjepa2.1-vit-base-384, you can try vjepa2.1-vit-base-384 online for free by clicking the link below.

apiantonio vjepa2.1-vit-base-384 online free url in huggingface.co:

https://huggingface.co/apiantonio/vjepa2.1-vit-base-384

vjepa2.1-vit-base-384 install

vjepa2.1-vit-base-384 is an open source model from GitHub that offers a free installation service, and any user can find vjepa2.1-vit-base-384 on GitHub to install. At the same time, huggingface.co provides the effect of vjepa2.1-vit-base-384 install, users can directly use vjepa2.1-vit-base-384 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

vjepa2.1-vit-base-384 install url in huggingface.co:

https://huggingface.co/apiantonio/vjepa2.1-vit-base-384

Url of vjepa2.1-vit-base-384

vjepa2.1-vit-base-384 huggingface.co Url

Provider of vjepa2.1-vit-base-384 huggingface.co

apiantonio
ORGANIZATIONS

Other API from apiantonio