FineWeb2 embeddings Dataset summary FineWeb2 embeddings is an extension of the FineWeb2 dataset, annotated with document level Snowflake's Arctic embed m v2.0 embeddings for 36 languages , making the dataset useful for a variety of tasks , including document clustering, filtering, and other multilingual research. Snowflake arctic embed m v2.0 has a sequence length limit of 8192 tokens, each document's embeddings are obtained by using the CLS token to embed each document. The embeddings were computed as part of our 🦊 JQL: Judging Quality across Languages project and will be the basis for an upcoming high quality subset of FineWeb2. We believe that they can be useful for other multilingual research and applications. For more details, see our paper Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models. Languages and subsets Subset name Language name Number of documents Disk size als Latn Tosk Albanian 8,597,826 18.18GB als Latn removed Tosk Albanian 4,055,619 12.60GB bul Cyrl Bulgarian 25,994,731 145.75GB bul Cyrl removed Bulgarian 31,046,392 122.45GB cat Latn Catalan 17,136,414 40.35GB cat Latn removed Catalan 20,738,135 41.77GB…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy