open mythos tiny test This model is a fine tuned version of on an unknown dataset. Model description More information needed Intended uses & limitations More information needed Training and evaluation data More information needed Training procedure Training hyperparameters The following hyperparameters were used during training: learning rate: 0.0003 train batch size: 4 eval batch size: 8 seed: 42 gradient accumulation steps: 2 total train batch size: 8 optimizer: Use OptimizerNames.ADAMW TORCH FUSED with betas=(0.9,0.95) and epsilon=1e 08 and optimizer args=No additional optimizer arguments lr scheduler type: cosine lr scheduler warmup steps: 500 training steps: 100 Training results Framework versions Transformers 5.6.0 Pytorch 2.11.0+cu130 Datasets 4.8.4 Tokenizers 0.22.2
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy