This dataset is uploaded in two places: here and additionally here as 'Aya Collection Language Split.' These datasets are identical in content but differ in structure of upload. This dataset is structured by folders split according to dataset name. The version here instead divides the Aya collection into folders split by language. We recommend you use the language split version if you are only interested in downloading data for a single or smaller set of languages, and this version if you want to download dataset according to data source or the entire collection. Dataset Summary The Aya Collection is a massive multilingual collection consisting of 513 million instances of prompts and completions covering a wide range of tasks. This collection incorporates instruction style templates from fluent speakers and applies them to a curated list of datasets, as well as translations of instruction style datasets into 101 languages. Aya Dataset, a human curated multilingual instruction and response dataset, is also part of this collection. See our paper for more details regarding the collection. Curated by: Contributors of Aya Open Science Intiative Language(s): 115 languages License: Apache…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy