Glot500 Corpus A dataset of natural language data collected by putting together more than 150 existing mono lingual and multilingual datasets together and crawling known multilingual websites. The focus of this dataset is on 500 extremely low resource languages. (More Languages still to be uploaded here) This dataset is used to train the Glot500 model. Homepage: homepage Repository: github Paper: acl, arxiv This dataset has the identical data format as the Taxi1500 Raw Data dataset, so that both datasets can be used in parallel seamlessly. Parts of the original Glot500 dataset cannot be published publicly. Please fill out [thi form]{https://docs.google.com/forms/d/1FHto 4wWYvEF3lz7DDo3P8wQqfS3WhpYfAu5vM95 qU/viewform?edit requested=true} to get access to these parts. Usage Replace nbl Latn with your specific language. Click to show supported languages: License We don't own any part of the data. The original source of each sentence of the data is indicated in dataset field. To see the copyright license of the original datasets visit here. We license the actual packaging, the metadata and the annotations of these data under the cc0 1.0. If you are a website/dataset owner and do not w…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy