Model Card for CLIP convnext base w 320.laion2B s13B b82K Table of Contents 1. Model Details 2. Uses 3. Training Details 4. Evaluation 5. Acknowledgements 6. Citation Model Details Model Description A series of CLIP ConvNeXt Base (w/ wide embed dim) models trained on subsets LAION 5B (https://laion.ai/blog/laion 5b/) using OpenCLIP (https://github.com/mlfoundations/open clip). Goals: Explore an alternative to ViT and ResNet (w/ AttentionPooling) CLIP models that scales well with model size and image resolution Firsts: First known ConvNeXt CLIP models trained at scale in the range of CLIP ViT B/16 and RN50x4 models First released model weights exploring increase of augmentation + regularization for image tower via adding (greater scale range of RRC, random erasing, stochastic depth) The models utilize the timm ConvNeXt Base model ( convnext base ) as the image tower, and the same text tower as the RN50x4 (depth 12, embed dim 640) model from OpenAI CLIP. The base models are trained at 256x256 image resolution and roughly match the RN50x4 models on FLOPs and activation counts. The models with 320 in the name are trained at 320x320. All models in this series were trained for 13B sample…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy