transformers ud japanese electra ginza 510 (sudachitra wordpiece, mC4 Japanese) This is an ELECTRA model pretrained on approximately 200M Japanese sentences extracted from the mC4 and finetuned by spaCy v3 on UD\ Japanese\ BCCWJ r2.8. The base pretrain model is megagonlabs/transformers ud japanese electra base discrimininator. The entire spaCy v3 model is distributed as a python package named ja ginza electra from PyPI along with GiNZA v5 which provides some custom pipeline components to recognize the Japanese bunsetu phrase structures. Try running it as below: Licenses The models are distributed under the terms of the MIT License. Acknowledgments This model is permitted to be published under the MIT License under a joint research agreement between NINJAL (National Institute for Japanese Language and Linguistics) and Megagon Labs Tokyo. Citations mC4 Contains information from mC4 which is made available under the ODC Attribution License. UD\ Japanese\ BCCWJ r2.8 GSK2014 A(2019)
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy