Model Description Language: Norwegian Bokmål and Nynorsk Developed by: HPLT Paper: arxiv.org/abs/2511.01066 Evaluation results: hf.co/datasets/HPLT/2508 datasets evals using HPLT E License: Apache 2.0 The HPLT's Llama 2b collection comprises monolingual decoder only language models pretrained by the HPLT team as part of the third release. The models are released as artifacts of our ablation studies on evaluating different corpora and sampling strategies across multiple languages: ⚖️ HPLT Pre 3.0 Comparison : Comparison of data deduplication strategies on a pre release version of HPLT 3.0 across nine selected languages (HPLT 3.0 pre release). 📚 Corpora Comparison : Evaluation of HPLT 2.0, HPLT 3.0, FineWeb 2.1.0, and MADLAD 400 1.0 on nine selected languages (HPLT 3.0 release). 🧰 Web Document Scorer (WDS) Comparison : Analysis of HPLT 3.0 corpora sampled using different WDS thresholds, focusing on Spanish and French (HPLT 3.0 release). Please find more details in our GitHub repository and pre print. Model Architecture All models follow the Llama architecture with 24 layers, 32 attention heads, and a sequence length of 2048. The tokenizer is Gemma 3 with the vocabulary size of 262K…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy