FineWeb2 embedded Dataset summary FineWeb2 embedded is an extension of the FineWeb2 dataset, annotated with document level XLM RoBERTa embeddings for 20 languages , making the dataset useful for a variety of tasks , including document clustering, filtering, and other multilingual research. Since XLM RoBERTa has a sequence length limit of 512 tokens, each document's embeddings are obtained by mean pooling 512 token chunks of the XLM RoBERTa output . Therefore, longer texts have more embeddings available (one per 512 tokens). The embeddings were initially computed as part of our FineWeb2 HQ dataset (a high quality subset of FineWeb2). However, we believe that they can be useful for other multilingual research and applications. For more details, see our paper Enhancing Multilingual LLM Pretraining with Model Based Data Selection. Languages and subsets Subset name Language name Number of documents Disk size : : rus Cyrl Russian 605,468,615 5.3T cmn Hani Chinese 578,332,129 4.4T deu Latn German 427,700,394 2.5T spa Latn Spanish 405,634,303 2.3T jpn Jpan Japanese 376,134,745 2.4T fra Latn French 332,646,715 2.0T ita Latn Italian 219,117,921 1.3T por Latn Portuguese 189,851,449 1.1T pol L…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy