KoRean based ELECTRA (KR ELECTRA) This is a release of a Korean specific ELECTRA model with comparable or better performances developed by the Computational Linguistics Lab at Seoul National University. Our model shows remarkable performances on tasks related to informal texts such as review documents, while still showing comparable results on other kinds of tasks. Released Model We pre trained our KR ELECTRA model following a base scale model of ELECTRA. We trained the model based on Tensorflow v1 using a v3 8 TPU of Google Cloud Platform. Model Details We followed the training parameters of the base scale model of ELECTRA. Hyperparameters model of layers embedding size hidden size of heads : : : : : Discriminator 12 768 768 12 Generator 12 768 256 4 Pretraining batch size train steps learning rates max sequence length generator size : : : : : 256 700000 2e 4 128 0.33333 Training Dataset 34GB Korean texts including Wikipedia documents, news articles, legal texts, news comments, product reviews, and so on. These texts are balanced, consisting of the same ratios of written and spoken data. Vocabulary vocab size 30,000 We used morpheme based unit tokens for our vocabulary based on th…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy