MMLW retrieval roberta large v2 MMLW (muszę mieć lepszą wiadomość) are neural text encoders for Polish. The second version is based on the same foundational model (polish roberta large v2), but the training process incorporated modern LLM based English retrievers and rerankers, which led to improved results. This model is optimized for information retrieval tasks. It can transform queries and passages to 1024 dimensional vectors. The model was developed using a two step procedure: In the first step, it was initialized with Polish RoBERTa checkpoint, and then trained with multilingual knowledge distillation method on a diverse corpus of 20 million Polish English text pairs. We utilised stella en 1.5B v5 as the teacher models for distillation. The second step involved fine tuning the model with contrastrive loss using a dataset consisting of over 4 million queries. Positive and negative passages for each query have been selected with the help of BAAI/bge reranker v2.5 gemma2 lightweight reranker. Usage (Sentence Transformers) The model supports both information retrieval and semantic textual similarity. For retrieval, queries should be prefixed with "[query]: " . For symmetric tasks…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy