Dataset Summary The VeraCruz Dataset is a comprehensive collection of Portuguese language content, showcasing the linguistic and cultural diversity of of Portuguese speaking regions. It includes around 190 million samples, organized by regional origin as indicated by URL metadata into primary categories. The primary categories are: Portugal (PT) : Samples with content URLs indicating a clear Portuguese origin. Brazil (BR) : Samples with content URLs indicating a clear Brazilian origin. Other : Samples where the URL metadata does not clearly indicate a Portuguese or Brazilian origin. These samples were further classified into "PT" or "BR" categories using the PeroVaz PT BR Classifier, which is trained specifically to distinguish between the European and Brazilian variations of Portuguese. Each entry in this category is supplemented with two extra columns: 'label' and 'score'. The 'label' column indicates the predicted category (PT or BR), and the 'score' column represents the probability of the predicted label. Source Data The VeraCruz Dataset is derived from the MyCulturaX dataset's Portuguese language segment, a comprehensive collection known for its broad linguistic coverage acro…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy