rugpt3small\ based\ on\ gpt2 The model architecture design, pretraining, and evaluation are documented in our preprint: A Family of Pretrained Transformer Language Models for Russian . The model was pretrained with sequence length 1024 using transformers by the SberDevices team on 80B tokens around 3 epochs. After that, the model was finetuned with the context size of 2048. Total training time took around one week on 32 GPUs. Authors + NLP core team RnD Telegram channel: + Dmitry Zmitrovich Cite us
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy