Dataset Card for OPUS 100 Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://opus.nlpl.eu/OPUS 100 Repository: https://github.com/EdinburghNLP/opus 100 corpus Paper: https://arxiv.org/abs/2004.11867 Paper: https://aclanthology.org/L10 1473/ Leaderboard: More Information Needed Point of Contact: More Information Needed Dataset Summary OPUS 100 is an English centric multilingual corpus covering 100 languages. OPUS 100 is English centric, meaning that all training pairs include English on either the source or target side. The corpus covers 100 languages (including English). The languages were selected based on the volume of parallel data available in OPUS. Supported Tasks and Leaderboards Translation. Languages OPUS 100 contains approximately 55M sentence…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy