Marqo FashionCLIP Model Card Marqo FashionCLIP and Marqo FashionSigLIP outperform the previous state of the art fashion CLIP models (see results below). Marqo FashionCLIP leverages Generalised Contrastive Learning (GCL) which allows the model to be trained on not just text descriptions but also categories, style, colors, materials, keywords and fine details to provide highly relevant search results on fashion products. The model was fine tuned from ViT B 16 (laion2b s34b b88k). Github Page : Marqo FashionCLIP Blog : Marqo Blog Usage Hugging Face The model can be loaded with AutoModel by OpenCLIP The model can be seamlessly used with OpenCLIP by Transformers.js You can also run the model in JavaScript with the Transformers.js library. First, install it from NPM using: Then, compute embeddings as follows: Benchmark Results Average evaluation results on 6 public multimodal fashion datasets (Atlas, DeepFashion (In shop), DeepFashion (Multimodal), Fashion200k, KAGL, and Polyvore) are reported below: Text To Image (Averaged across 6 datasets) Model AvgRecall Recall@1 Recall@10 MRR Marqo FashionCLIP 0.192 0.094 0.290 0.200 FashionCLIP2.0 0.163 0.077 0.249 0.165 OpenFashionCLIP 0.132 0.060…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy