Model card for PE Core L 14 336 This is an OpenCLIP (image + text) remaped version of the the original [\[📃 Tech Report\]](https://arxiv.org/abs/2504.13181) [\[📂 PE Github (original weights)\]](https://github.com/facebookresearch/perception models/) [\[📂 OpenCLIP Github (these weights)\]](https://github.com/mlfoundations/open clip) Perception Encoder (PE) is a state of the art encoder for image and video understanding trained via simple vision language learning. It was introduced in "Perception Encoder: The best visual embeddings are not at the output of the network". Model Developer : Meta Model Overview : Perception Encoder (PE) is a family of large scale vision encoder models with state of the art performance on a large variety of vision tasks. By using a robust contrastive pretraining recipe and finetuning on synthetically aligned videos, PE not only outperforms all existing models on classification and retrieval, but it also internally produces strong, general features that scale for downstream tasks. PE unlocks the ability for large scale contrastive pretraining to transfer to downstream tasks with alignment tuning to capitalize on those general features. Perception Encode…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy