Qwen3 4B SFT science 1e 5 This model is a fine tuned version of Qwen/Qwen3 4B on the dolci science train dataset. It achieves the following results on the evaluation set: Loss: 0.6778 Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning rate: 1e 05 train batch size: 2 eval batch size: 2 seed: 42 distributed type: multi GPU num devices: 4 gradient accumulation steps: 16 total train batch size: 128 total eval batch size: 8 optimizer: Use OptimizerNames.ADAMW TORCH FUSED with betas=(0.9,0.999) and epsilon=1e 08 and optimizer args=No additional optimizer arguments lr scheduler type: cosine lr scheduler warmup steps: 0.05 num epochs: 3.0 Training results Training Loss Epoch Step Validation Loss : : : : : : : : 0.8065 0.2985 230 0.7250 0.6763 0.5969 460 0.7040 0.7030 0.8954 690 0.6914 0.6122 1.1933 920 0.6877 0.6361 1.4918 1150 0.6827 0.6499 1.7903 1380 0.6778 0.5879 2.0882 1610 0.6838 0.5390 2.3867 1840 0.6826 0.6058 2.6852 2070 0.6820 0.6097 2.9836 2300 0.6816 Framework versions Transf…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy