Dataset A manually annotated dataset was created to enable supervised training for the FinBERT PT BR model, which focuses on sentiment analysis of Brazilian Portuguese financial texts. More than 1.4 million financial news texts in Portuguese were collected and used for the initial language modeling phase. From this corpus, a sample of 1,000 texts was manually annotated with sentiment labels. Annotation Process Three annotators participated in the process. All texts were annotated by at least two different annotators. The defined categories were: POSITIVE, NEGATIVE, NEUTRAL, and NOT APPLICABLE. Annotation guideline: "Classify the text based on whether it would imply a Positive, Negative, or Neutral return. Use 'Not applicable' for texts unrelated to finance, involving politics, or nonsensical content." After a calibration step to ensure agreement among annotators, the full sample was annotated. Final Dataset Composition Out of the 1,000 annotated texts: 497 texts were discarded due to lack of agreement or being labeled as “Not applicable”. The final training dataset contains 503 texts: 160 positive 203 negative 140 neutral Agreement Metrics Agreement rate: 90.4% Krippendorff’s alpha…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy