ParsBERT (v2.0) A Transformer based Model for Persian Language Understanding We reconstructed the vocabulary and fine tuned the ParsBERT v1.1 on the new Persian corpora in order to provide some functionalities for using ParsBERT in other scopes! Please follow the ParsBERT repo for the latest information about previous and current models. Introduction ParsBERT is a monolingual language model based on Google’s BERT architecture. This model is pre trained on large Persian corpora with various writing styles from numerous subjects (e.g., scientific, novels, news) with more than 3.9M documents, 73M sentences, and 1.3B words. Paper presenting ParsBERT: arXiv:2005.12515 Intended uses & limitations You can use the raw model for either masked language modeling or next sentence prediction, but it's mostly intended to be fine tuned on a downstream task. See the model hub to look for fine tuned versions on a task that interests you. How to use TensorFlow 2.0 Pytorch Training ParsBERT trained on a massive amount of public corpora (Persian Wikidumps, MirasText) and six other manually crawled text data from a various type of websites (BigBang Page scientific , Chetor lifestyle , Eligasht itinerar…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy