library name: transformers license: other base model: Qwen/Qwen3 8B tags: llama factory full generated from trainer model index: name: Qwen3 feedback results: [] Qwen3 feedback This model is a fine tuned version of Qwen/Qwen3 8B on the sft dataset. It achieves the following results on the evaluation set: Loss: 0.0228 Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning rate: 1e 05 train batch size: 2 eval batch size: 2 seed: 42 distributed type: multi GPU num devices: 7 total train batch size: 14 total eval batch size: 14 optimizer: Use OptimizerNames.ADAMW TORCH with betas=(0.9,0.999) and epsilon=1e 08 and optimizer args=No additional optimizer arguments lr scheduler type: cosine lr scheduler warmup ratio: 0.05 num epochs: 1.0 mixed precision training: Native AMP Training results Training Loss Epoch Step Validation Loss : : : : : : : : 0.0367 0.0819 1000 0.0367 0.0274 0.1638 2000 0.0304 0.0257 0.2457 3000 0.0294 0.0242 0.3276 4000 0.0265 0.0239 0.4095 5000 0.0256 0.0235 0.4914 600…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy