robertuito base uncased RoBERTuito A pre trained language model for social media text in Spanish PAPER Github Repository RoBERTuito is a pre trained language model for user generated content in Spanish, trained following RoBERTa guidelines on 500 million tweets. RoBERTuito comes in 3 flavors: cased, uncased, and uncased+deaccented. We tested RoBERTuito on a benchmark of tasks involving user generated text in Spanish. It outperforms other pre trained language models for this language such as BETO , BERTin and RoBERTa BNE . The 4 tasks selected for evaluation were: Hate Speech Detection (using SemEval 2019 Task 5, HatEval dataset), Sentiment and Emotion Analysis (using TASS 2020 datasets), and Irony detection (using IrosVa 2019 dataset). model hate speech sentiment analysis emotion analysis irony detection score : : : : : : robertuito uncased 0.801 ± 0.010 0.707 ± 0.004 0.551 ± 0.011 0.736 ± 0.008 0.6987 robertuito deacc 0.798 ± 0.008 0.702 ± 0.004 0.543 ± 0.015 0.740 ± 0.006 0.6958 robertuito cased 0.790 ± 0.012 0.701 ± 0.012 0.519 ± 0.032 0.719 ± 0.023 0.6822 roberta bne 0.766 ± 0.015 0.669 ± 0.006 0.533 ± 0.011 0.723 ± 0.017 0.6726 bertin 0.767 ± 0.005 0.665 ± 0.003 0.518 ± 0.012…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy