GottBERT: A pure German language model GottBERT is the first German only RoBERTa model, pre trained on the German portion of the first released OSCAR dataset. This model aims to provide enhanced natural language processing (NLP) performance for the German language across various tasks, including Named Entity Recognition (NER), text classification, and natural language inference (NLI). GottBERT has been developed in two versions: a base model and a large model , tailored specifically for German language tasks. Model Type : RoBERTa Language : German Base Model : 12 layers, 125 million parameters Large Model : 24 layers, 355 million parameters License : MIT This was presented in GottBERT: a pure German Language Model. Pretraining Details Corpus : German portion of the OSCAR dataset (Common Crawl). Data Size : Unfiltered: 145GB (~459 million documents) Filtered: 121GB (~382 million documents) Preprocessing : Filtering included correcting encoding errors (e.g., erroneous umlauts), removing spam and non German documents using language detection and syntactic filtering. Filtering Metrics Stopword Ratio : Detects spam and meaningless content. Punctuation Ratio : Detects abnormal punctuatio…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy