UmBERTo Commoncrawl Cased UmBERTo is a Roberta based Language Model trained on large Italian Corpora and uses two innovative approaches: SentencePiece and Whole Word Masking. Now available at github.com/huggingface/transformers Marco Lodola, Monument to Umberto Eco, Alessandria 2019 Dataset UmBERTo Commoncrawl Cased utilizes the Italian subcorpus of OSCAR as training set of the language model. We used deduplicated version of the Italian corpus that consists in 70 GB of plain text data, 210M sentences with 11B words where the sentences have been filtered and shuffled at line level in order to be used for NLP research. Pre trained model Model WWM Cased Tokenizer Vocab Size Train Steps Download umberto commoncrawl cased v1 YES YES SPM 32K 125k Link This model was trained with SentencePiece and Whole Word Masking. Downstream Tasks These results refers to umberto commoncrawl cased model. All details are at Umberto Official Page. Named Entity Recognition (NER) Dataset F1 Precision Recall Accuracy ICAB EvalITA07 87.565 86.596 88.556 98.690 WikiNER ITA 92.531 92.509 92.553 99.136 Part of Speech (POS) Dataset F1 Precision Recall Accuracy UD Italian ISDT 98.870 98.861 98.879 98.977 UD Italia…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy