rugpt3large\ based\ on\ gpt2 The model architecture design, pretraining, and evaluation are documented in our preprint: A Family of Pretrained Transformer Language Models for Russian . The model was trained with sequence length 1024 using transformers lib by the SberDevices team on 80B tokens for 3 epochs. After that, the model was finetuned 1 epoch with sequence length 2048. Total training time was around 14 days on 128 GPUs for 1024 context and a few days on 16 GPUs for 2048 context. The final perplexity on the test set is 13.6 . Authors + NLP core team RnD Telegram channel: + Dmitry Zmitrovich Cite us
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy