gte multilingual mlm base We introduce mGTE series, new generalized text encoder, embedding and reranking models that support 75 languages and the context length of up to 8192. The models are built upon the transformer++ encoder backbone (BERT + RoPE + GLU, code refer to Alibaba NLP/new impl) as well as the vocabulary of XLM R . This text encoder ( mGTE MLM 8192 in our paper) outperforms the same sized previous state of the art XLM R base in both GLUE and XTREME R. Developed by : Institute for Intelligent Computing, Alibaba Group Model type : Text Encoder Paper : mGTE: Generalized Long Context Text Representation and Reranking Models for Multilingual Text Retrieval. Model list Models Language Model Size Max Seq. Length GLUE XTREME R : : : : : : : : : : : : gte multilingual mlm base Multiple 306M 8192 83.47 64.44 gte en mlm base English 8192 85.61 gte en mlm large English 8192 87.58 Training Details Training Data Masked language modeling (MLM): c4 en , mc4 , skypile , Wikipedia , CulturaX , etc (refer to paper appendix A.1) Training Procedure To enable the backbone model to support a context length of 8192, we adopted a multi stage training strategy. The model first undergoes prelim…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy