Dataset Card for xP3 Table of Contents Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Additional Information Licensing Information Citation Information Contributions Dataset Description Repository: https://github.com/bigscience workshop/xmtf Paper: Crosslingual Generalization through Multitask Finetuning Point of Contact: Niklas Muennighoff Dataset Summary xP3 (Crosslingual Public Pool of Prompts) is a collection of prompts & datasets across 46 of languages & 16 NLP tasks. It is used for the training of BLOOMZ and mT0, multilingual language models capable of following human instructions in dozens of languages zero shot. Creation: The dataset can be recreated using instructions available here. We provide this version to save processing time and ease reproducibility. Languages: 46 (Can be extended by recreating with more splits) xP3 Dataset Family: Name Explanation Example models xP3x Mixture of 17 tasks in 277 languages with English prompts WIP Join us at Project Aya @ C4AI to help! xP3 Mixture of 13 training tasks in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy