Multilingual E5 large instruct Multilingual E5 Text Embeddings: A Technical Report. Liang Wang, Nan Yang, Xiaolong Huang, Linjun Yang, Rangan Majumder, Furu Wei, arXiv 2024 This model has 24 layers and the embedding size is 1024. Usage Below are examples to encode queries and passages from the MS MARCO passage ranking dataset. Transformers Sentence Transformers Infinity Usage with Infinity: Supported Languages This model is initialized from xlm roberta large and continually trained on a mixture of multilingual datasets. It supports 100 languages from xlm roberta, but low resource languages may see performance degradation. Training Details Initialization : xlm roberta large First stage : contrastive pre training with 1 billion weakly supervised text pairs. Second stage : fine tuning on datasets from the E5 mistral paper. MTEB Benchmark Evaluation Check out unilm/e5 to reproduce evaluation results on the BEIR and MTEB benchmark. FAQ 1. Do I need to add instructions to the query? Yes, this is how the model is trained, otherwise you will see a performance degradation. The task definition should be a one sentence instruction that describes the task. This is a way to customize text embed…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy