Dataset Card for Parallel Sentences CCMatrix This dataset contains parallel sentences (i.e. English sentence + the same sentences in another language) for numerous other languages. The texts originate from the CCMatrix dataset. Related Datasets The following datasets are also a part of the Parallel Sentences collection: parallel sentences europarl parallel sentences global voices parallel sentences muse parallel sentences jw300 parallel sentences news commentary parallel sentences opensubtitles parallel sentences talks parallel sentences tatoeba parallel sentences wikimatrix parallel sentences wikititles parallel sentences ccmatrix These datasets can be used to train multilingual sentence embedding models. For more information, see sbert.net Multilingual Models. Dataset Subsets en ... subsets Columns: "english", "non english" Column types: str , str Examples: Collection strategy: Processing the data from yhavinga/ccmatrix and reformatting it in Parquet and with "english" and "non english" columns. Deduplified: No
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy