lua mistral 2L tiny This model is a fine tuned version of on the nilq/small lua stack dataset. It achieves the following results on the evaluation set: Loss: 1.6229 Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning rate: 0.0006 train batch size: 64 eval batch size: 8 seed: 42 optimizer: Adam with betas=(0.9,0.999) and epsilon=1e 08 lr scheduler type: cosine num epochs: 3.0 Training results Framework versions Transformers 4.38.1 Pytorch 2.2.0+cu121 Datasets 2.17.1 Tokenizers 0.15.2
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy