gte base General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi stage Contrastive Learning The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer three different sizes of models, including GTE large, GTE base, and GTE small. The GTE models are trained on a large scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval , semantic textual similarity , text reranking , etc. Metrics We compared the performance of the GTE models with other popular text embedding models on the MTEB benchmark. For more detailed comparison results, please refer to the MTEB leaderboard. Model Name Model Size (GB) Dimension Sequence Length Average (56) Clustering (11) Pair Classification (3) Reranking (4) Retrieval (15) STS (10) Summarization (1) Classification (12) : : : : : : : : : : : : : : : : : : : : : : : : gte large 0.67 1024 512 63.13 46.84 85.00 59.13 52.22 83.35 31.66 73.33 gte base 0.22 768 512 62.39 46.2 84.57 58.61 51.14 82.3 31.17 73.01 e5 large v2 1.34 1024 512…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy