ruBert base The model architecture design, pretraining, and evaluation are documented in our preprint: A Family of Pretrained Transformer Language Models for Russian . The model is pretrained by the SberDevices team. Task: mask filling Type: encoder Tokenizer: BPE Dict size: 120 138 Num Parameters: 178 M Training Data Volume 30 GB Authors + NLP core team RnD Telegram channel: + Dmitry Zmitrovich Cite us
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy