gte large zh General Text Embeddings (GTE) model. Towards General Text Embeddings with Multi stage Contrastive Learning The GTE models are trained by Alibaba DAMO Academy. They are mainly based on the BERT framework and currently offer different sizes of models for both Chinese and English Languages. The GTE models are trained on a large scale corpus of relevance text pairs, covering a wide range of domains and scenarios. This enables the GTE models to be applied to various downstream tasks of text embeddings, including information retrieval , semantic textual similarity , text reranking , etc. Model List Models Language Max Sequence Length Dimension Model Size : : : : : : : : : : GTE large zh Chinese 512 1024 0.67GB GTE base zh Chinese 512 512 0.21GB GTE small zh Chinese 512 512 0.10GB GTE large English 512 1024 0.67GB GTE base English 512 512 0.21GB GTE small English 512 384 0.10GB Metrics We compared the performance of the GTE models with other popular text embedding models on the MTEB (CMTEB for Chinese language) benchmark. For more detailed comparison results, please refer to the MTEB leaderboard. Evaluation results on CMTEB Model Model Size (GB) Embedding Dimensions Sequence…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy