WMT24++ This repository contains the human translation and post edit data for the 55 en xx language pairs released in the publication WMT24++: Expanding the Language Coverage of WMT24 to 55 Languages & Dialects. If you are interested in the MT/LLM system outputs and automatic metric scores, please see MTME. If you are interested in the images of the source URLs for each document, please see here. Schema Each language pair is stored in its own jsonl file. Each row is a serialized JSON object with the following fields: lp : The language pair (e.g., "en de DE" ). domain : The domain of the source, either "canary" , "news" , "social" , "speech" , or "literary" . document id : The unique ID that identifies the document the source came from. segment id : The globally unique ID that identifies the segment. is bad source : A Boolean that indicates whether this source is low quality (e.g., HTML, URLs, emoijs). In the paper, the segments marked as true were removed from the evaluation, and we recommend doing the same. source : The English source text. target : The post edit of original target . We recommend using the post edit as the default reference. original target : The original referenc…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy