Model Overview
: Perception Encoder (PE) is a family of large-scale vision encoder models with state-of-the-art performance on a large variety of vision tasks. By using a robust contrastive pretraining recipe and finetuning on synthetically aligned videos, PE not only outperforms all existing models on classification and retrieval, but it also internally produces strong, general features that scale for downstream tasks. PE unlocks the ability for large-scale contrastive pretraining to transfer to downstream tasks with alignment tuning to capitalize on those general features.
Perception Encoder: Core
PE core is our base model trained with our robust image pretraining schedule and finetuned on the data generated by our synthetic video data engine.
Model Configurations
PE core curently comes in 3 sizes. PE core G is the main checkpoint, with L and B models distilled from it.
Scale
Tower
Params
Width
Depth
MLP
Heads
CLIP Dim
Resolution / Context Len
B/16
Vision
0.09B
768
12
3072
12
1024
224px
Text
0.31B
1024
24
4096
16
1024
32 tokens
L/14
Vision
0.32B
1024
24
4096
16
1024
336px
Text
0.31B
1024
24
4096
16
1024
32 tokens
G/14
Vision
1.88B
1536
50
8960
16
1280
448px
Text
0.47B
1280
24
5120
20
1280
72 tokens
All PE core models use an attention pooling block with 8 heads on top of the vision tower. The L and B models
additionally
have a class token for global aggregation. See the paper for more details.
Model Performance
PE core obtains extremely strong results across the board on zero-shot image classification and retrieval
as well as
zero-shot video classification and retrieval. We present a sample of its performance across those domains below.
If you find our code useful for your research, please consider citing:
@article{bolya2025PerceptionEncoder,
title={Perception Encoder: The best visual embeddings are not at the output of the network},
author={Daniel Bolya and Po-Yao Huang and Peize Sun and Jang Hyun Cho and Andrea Madotto and Chen Wei and Tengyu Ma and Jiale Zhi and Jathushan Rajasegaran and Hanoona Rasheed and Junke Wang and Marco Monteiro and Hu Xu and Shiyu Dong and Nikhila Ravi and Daniel Li and Piotr Doll{\'a}r and Christoph Feichtenhofer},
journal={arXiv},
year={2025}
}
@article{cho2025PerceptionLM,
title={PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding},
author={Jang Hyun Cho and Andrea Madotto and Effrosyni Mavroudi and Triantafyllos Afouras and Tushar Nagarajan and Muhammad Maaz and Yale Song and Tengyu Ma and Shuming Hu and Hanoona Rasheed and Peize Sun and Po-Yao Huang and Daniel Bolya and Suyog Jain and Miguel Martin and Huiyu Wang and Nikhila Ravi and Shashank Jain and Temmy Stark and Shane Moon and Babak Damavandi and Vivian Lee and Andrew Westbury and Salman Khan and Philipp Kr\"{a}henb\"{u}hl and Piotr Doll{\'a}r and Lorenzo Torresani and Kristen Grauman and Christoph Feichtenhofer},
journal={arXiv},
year={2025}
}
Runs of facebook PE-Core-L14-336 on huggingface.co
442.4K
Total runs
0
24-hour runs
12.1K
3-day runs
33.5K
7-day runs
-34.5K
30-day runs
More Information About PE-Core-L14-336 huggingface.co Model
PE-Core-L14-336 huggingface.co is an AI model on huggingface.co that provides PE-Core-L14-336's model effect (), which can be used instantly with this facebook PE-Core-L14-336 model. huggingface.co supports a free trial of the PE-Core-L14-336 model, and also provides paid use of the PE-Core-L14-336. Support call PE-Core-L14-336 model through api, including Node.js, Python, http.
PE-Core-L14-336 huggingface.co is an online trial and call api platform, which integrates PE-Core-L14-336's modeling effects, including api services, and provides a free online trial of PE-Core-L14-336, you can try PE-Core-L14-336 online for free by clicking the link below.
facebook PE-Core-L14-336 online free url in huggingface.co:
PE-Core-L14-336 is an open source model from GitHub that offers a free installation service, and any user can find PE-Core-L14-336 on GitHub to install. At the same time, huggingface.co provides the effect of PE-Core-L14-336 install, users can directly use PE-Core-L14-336 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.