Model Card for ru en RoSBERTa The ru en RoSBERTa is a general text embedding model for Russian. The model is based on ruRoBERTa and fine tuned with ~4M pairs of supervised, synthetic and unsupervised data in Russian and English. Tokenizer supports some English tokens from RoBERTa tokenizer. For more model details please refer to our article. Usage The model can be used as is with prefixes. It is recommended to use CLS pooling. The choice of prefix and pooling depends on the task. We use the following basic rules to choose a prefix: "search query: " and "search document: " prefixes are for answer or relevant paragraph retrieval "classification: " prefix is for symmetric paraphrasing related tasks (STS, NLI, Bitext Mining) "clustering: " prefix is for any tasks that rely on thematic features (topic classification, title body retrieval) To better tailor the model to your needs, you can fine tune it with relevant high quality Russian and English datasets. Below are examples of texts encoding using the Transformers and SentenceTransformers libraries. Transformers SentenceTransformers or using prompts (sentence transformers =2.4.0): Citation Limitations The model is designed to process t…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy