This is an updated version of cointegrated/rubert tiny: a small Russian BERT based encoder with high quality sentence embeddings. This post in Russian gives more details. The differences from the previous version include: a larger vocabulary: 83828 tokens instead of 29564; larger supported sequences: 2048 instead of 512; sentence embeddings approximate LaBSE closer than before; meaningful segment embeddings (tuned on the NLI task) the model is focused only on Russian. The model should be used as is to produce sentence embeddings (e.g. for KNN classification of short texts) or fine tuned for a downstream task. Sentence embeddings can be produced as follows: Alternatively, you can use the model with sentence transformers : For those who want to run the inference with VLLM, there is a vLLM optimized version of this model: WpythonW/rubert tiny2 vllm
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy