zephyr 7b sft full This model is a fine tuned version of mistralai/Mistral 7B v0.1 on the HuggingFaceH4/ultrachat 200k dataset. It achieves the following results on the evaluation set: Loss: 0.9353 Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning rate: 2e 05 train batch size: 16 eval batch size: 8 seed: 42 distributed type: multi GPU num devices: 8 total train batch size: 128 total eval batch size: 64 optimizer: Adam with betas=(0.9,0.999) and epsilon=1e 08 lr scheduler type: cosine lr scheduler warmup ratio: 0.1 num epochs: 1 Training results Training Loss Epoch Step Validation Loss : : : : : : : : 0.9075 1.0 1090 0.9353 Framework versions Transformers 4.36.2 Pytorch 2.1.2+cu121 Datasets 2.14.6 Tokenizers 0.15.0
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy