Dataset Card for DBpedia14 Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: More Information Needed Repository: https://github.com/zhangxiangxiao/Crepe Paper: https://arxiv.org/abs/1509.01626 Point of Contact: Xiang Zhang Dataset Summary The DBpedia ontology classification dataset is constructed by picking 14 non overlapping classes from DBpedia 2014. They are listed in classes.txt. From each of thse 14 ontology classes, we randomly choose 40,000 training samples and 5,000 testing samples. Therefore, the total size of the training dataset is 560,000 and testing dataset 70,000. There are 3 columns in the dataset (same for train and test splits), corresponding to class index (1 to 14), title and content. The title and content are escaped using double quotes (")…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy