Dataset Card for Flores200 Table of Contents Dataset Card for Flores200 Table of Contents Dataset Description Dataset Summary Supported Tasks and Leaderboards Languages Dataset Structure Data Instances Data Fields Data Splits Dataset Creation Additional Information Dataset Curators Licensing Information Citation Information Dataset Description Home: Flores Repository: Github Dataset Summary FLORES is a benchmark dataset for machine translation between English and low resource languages. The creation of FLORES200 doubles the existing language coverage of FLORES 101. Given the nature of the new languages, which have less standardization and require more specialized professional translations, the verification process became more complex. This required modifications to the translation workflow. FLORES 200 has several languages which were not translated from English. Specifically, several languages were translated from Spanish, French, Russian and Modern Standard Arabic. Moreover, FLORES 200 also includes two script alternatives for four languages. FLORES 200 consists of translations from 842 distinct web articles, totaling 3001 sentences. These sentences are divided into three splits:…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy