Slovak GPT J 405M Slovak GPT J 405M is the second model released in Slovak GPT J series after its smaller variant Slovak GPT J 162M. Since then a larger Slovak GPT J 1.4B was released. Model Description Model is based on GPT J and has over 405M trainable parameters. Hyperparameter Value \\(n {parameters}\\) 405,677,136 \\(n {layers}\\) 24 \\(d {model}\\) 1024 \\(d {ff}\\) 16384 \\(n {heads}\\) 16 \\(d {head}\\) 256 \\(n {ctx}\\) 2048 \\(n {vocab}\\) 50256 (same tokenizer as GPT 2/3†) Positional Encoding Rotary Position Embedding (RoPE) RoPE Dimensions 64 † ByteLevelBPETokenizer was trained on the same Slovak corpus. Training data Slovak GPT J models were trained on a privately collected dataset consisting of predominantly Slovak text spanning different categories, e.g. web, news articles or even biblical texts in total, over 40GB of text data was used to train this model. The dataset was preprocessed and cleaned in a specific way that involves minor but a few caveats, so in order to achieve the expected performance, feel free to refer to [How to use] section. Please, keep in mind that despite the effort to remove inappropriate corpus, the model still might generate se…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy