Dataset Card for [Dataset Name] Table of Contents Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Curation Rationale Source Data Annotations Personal and Sensitive Information Considerations for Using the Data Social Impact of Dataset Discussion of Biases Other Known Limitations Additional Information Dataset Curators Licensing Information Citation Information Contributions Dataset Description Homepage: https://github.com/Helsinki NLP/Tatoeba Challenge/ Repository: https://github.com/Helsinki NLP/Tatoeba Challenge/ Paper: The Tatoeba Translation Challenge – Realistic Data Sets for Low Resource and Multilingual MT Leaderboard: Point of Contact: Jörg Tiedemann Dataset Summary The Tatoeba Translation Challenge is a multilingual data set of machine translation benchmarks derived from user contributed translations collected by Tatoeba.org and provided as parallel corpus from OPUS. This dataset includes test and development data sorted by language pair. It includes test sets for hundreds of language pairs and is continuously updated. Please, check the version number…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy