fine tuned student kd rag transformer 20epochs This model was trained from scratch on the None dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning rate: 5e 05 train batch size: 8 eval batch size: 8 seed: 42 optimizer: Use adamw torch with betas=(0.9,0.999) and epsilon=1e 08 and optimizer args=No additional optimizer arguments lr scheduler type: linear num epochs: 5 mixed precision training: Native AMP Training results Framework versions Transformers 4.51.3 Pytorch 2.7.0+cu126 Datasets 3.6.0 Tokenizers 0.21.2
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy