Model Card for FRIDA FRIDA is a full scale finetuned general text embedding model inspired by denoising architecture based on T5. The model is based on the encoder part of FRED T5 model and continues research of text embedding models (ruMTEB, ru en RoSBERTa). It has been pre trained on a Russian English dataset and fine tuned for improved performance on the target task. For more model details please refer to our article (RU). Usage The model can be used as is with prefixes. It is recommended to use CLS pooling. The choice of prefix and pooling depends on the task. We use the following basic rules to choose a prefix: "search query: " and "search document: " prefixes are for answer or relevant paragraph retrieval "paraphrase: " prefix is for symmetric paraphrasing related tasks (STS, paraphrase mining, deduplication) "categorize: " prefix is for asymmetric matching of document title and body (e.g. news, scientific papers, social posts) "categorize sentiment: " prefix is for any tasks that rely on sentiment features (e.g. hate, toxic, emotion) "categorize topic: " prefix is intended for tasks where you need to group texts by topic "categorize entailment: " prefix is for textual entail…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy