HPLT v2.0 Cleaned French This is one of the decoder only language models trained on HPLT2.0 cleaned. All the HPLT decoder only models use the same hyper parameters, roughly following the llama architecture with 2.15B parameters in total: hidden size: 2048 attention heads: 32 layers: 24 sequence length: 2048 Intermediate checkpoints We are releasing intermediate checkpoints for each model at intervals of every 1000 training steps in separate branches. The naming convention is checkpoint 00xxxx00 : for example, checkpoint 0005000 . The checkpoints range from checkpoint 0001000 to checkpoint 0047684 and the latter is in the main branch. Cite us
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy