HPLT2 embeddings Dataset summary HPLT2 embeddings is an extension of the HPLT2 dataset, annotated with document level Snowflake's Arctic embed m v2.0 embeddings for 35 languages , making the dataset useful for a variety of tasks , including document clustering, filtering, and other multilingual research. Snowflake arctic embed m v2.0 has a sequence length limit of 8192 tokens, each document's embeddings are obtained by using the CLS token to embed each document. The embeddings were computed as part of our 🦊 JQL: Judging Quality across Languages project and will be the basis for an upcoming high quality subset of HPLT2. We believe that they can be useful for other multilingual research and applications. For more details, see our paper Judging Quality Across Languages: A Multilingual Approach to Pretraining Data Filtering with Language Models. Usage You can load the dataset in Python using e.g.pandas: Origin of the Dataset This dataset, derived from HPLT2, includes web content collected from 2013 to 2024. As HPLT2 is sourced from the broader internet, it may contain some personally identifiable information (PII), despite efforts to anonymize email addresses and public IP addresses d…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy