Model card for CLIP convnext large d 320.laion2B s29B b131K ft soup Table of Contents 1. Model Details 2. Uses 3. Training Details 4. Evaluation 5. Acknowledgements 6. Citation Model Details Model Description A series of CLIP ConvNeXt Large (w/ extra text depth, vision MLP head) models trained on the LAION 2B (english) subset of LAION 5B using OpenCLIP. The models utilize: the timm ConvNeXt Large model ( convnext large ) as the image tower a MLP ( fc gelu drop fc ) head in vision tower instead of the single projection of other CLIP models a text tower with same width but 4 layers more depth than ViT L / RN50x16 models (depth 16, embed dim 768). This 320x320 resolution model is a soup (weight average) of 3 fine tunes of CLIP convnext large d.laion2B s26B b102K augreg at a higher resolution. It is an average of 3 fine tunes from the final checkpoint of the original 256x256 training run w/ an additional ~2 3B samples for each fine tune and a lower learning rate. Each fine tune was a different learning rate (1e 4, 6e 5, 5e 5), and diff of samples (3.2B, 2B, 2.5B). At 320x320, the ConvNext Large D is significantly more efficient than the L/14 model at 336x336 that OpenAI fine tuned. L/1…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy