Dataset Card for xP3x Table of Contents Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Additional Information Licensing Information Citation Information Contributions Dataset Description Repository: https://github.com/bigscience workshop/xmtf Paper: Crosslingual Generalization through Multitask Finetuning Point of Contact: Niklas Muennighoff Dataset Summary xP3x (Crosslingual Public Pool of Prompts eXtended) is a collection of prompts & datasets across 277 languages & 16 NLP tasks. It contains all of xP3 + much more! It is used for training future contenders of mT0 & BLOOMZ at project Aya @Cohere Labs 🧡 Creation: The dataset can be recreated using instructions available here together with the file in this repository named xp3x create.py . We provide this version to save processing time. Languages: 277 xP3 Dataset Family: Name Explanation Example models xP3x Mixture of 17 tasks in 277 languages with English prompts WIP Join us at Project Aya @ C4AI to help! xP3 Mixture of 13 training tasks in 46 languages with English…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy