ruRoberta large The model architecture design, pretraining, and evaluation are documented in our preprint: A Family of Pretrained Transformer Language Models for Russian . The model is pretrained by the SberDevices team. Task: mask filling Type: encoder Tokenizer: BBPE Dict size: 50 257 Num Parameters: 355 M Training Data Volume 250 GB Authors + NLP core team RnD Telegram channel: + Dmitry Zmitrovich Cite us
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy