Model card for CLIP convnext xxlarge laion2B s34B b82K augreg Table of Contents 1. Model Details 2. Uses 3. Training Details 4. Evaluation 5. Acknowledgements 6. Citation Model Details Model Description A series of CLIP ConvNeXt XXLarge (a custom timm ConvNeXt size) models trained on LAION 2B (english), a subset of LAION 5B, using OpenCLIP. Model Dataset Resolution AugReg Top 1 ImageNet Zero Shot (%) convnext xxlarge.laion2b s34b b82k augreg LAION 2B 256x256 RRC (0.33, 1.0), RE (0.35), SD (0.1) 79.1 convnext xxlarge.laion2b s34b b82k augreg rewind LAION 2B 256x256 RRC (0.3, 1.0), RE (0.4), SD (0.1) 79.3 convnext xxlarge.laion2b s34b b82k augreg soup LAION 2B 256x256 N/A 79.4 RRC = Random Resize Crop (crop pcts), RE = Random Erasing (prob), SD = Stochastic Depth (prob) image tower only The core training run was performed in pieces over a period of ~ 2 months. The global batch size for the core run was 81920. The last ~10% of training was re done at a 95744 global batch size w/ higher LR and aug than original finish. The two were averaged together in a 'soup'. See more details in Training Details. Goals: Push the size of largest convolutional CLIP image tower into the performance ran…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy