HPLT2 Edu scores Dataset summary HPLT2 JQL Education is a model annotated language subset of HPLT2 , spanning 35 languages . Our model annotations allow for a filtering that achieves higher quality training outcomes without excessively aggressive data reduction. The original FW2 heuristic filtering method serves as our baseline, providing reference points for both the volume of retained tokens and downstream model performance. For example, in the Spanish language case, applying the 0.6 threshold retains over 9% more tokens than FW2 train while still surpassing its quality . HPLT2 Edu scores was created based on scores assigned by a deep learning classifier trained to identify educational samples using Snowflake's Arctic embed m v2.0 embeddings. For all training ablations, we used dense decoder only models with 2 billion parameters , following the LLaMA architecture. For more details, see our paper https://arxiv.org/abs/2505.22232. The approach as described in the paper is easy to extend to other languages as well, and we might consider adding new languages to an upcoming version of the present dataset. We also separately release the computed general purpose embedding vectors for th…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy