Google's T5 Version 1.1 LM Adapted Version 1.1 LM Adapted T5 Version 1.1 LM Adapted includes the following improvements compared to the original T5 model: GEGLU activation in feed forward hidden layer, rather than ReLU see here. Dropout was turned off in pre training (quality win). Dropout should be re enabled during fine tuning. Pre trained on C4 only without mixing in the downstream tasks. no parameter sharing between embedding and classifier layer "xl" and "xxl" replace "3B" and "11B". The model shapes are a bit different larger d model and smaller num heads and d ff . and is pretrained on both the denoising and language modeling objective. More specifically, this checkpoint is initialized from T5 Version 1.1 Small and then trained for an additional 100K steps on the LM objective discussed in the T5 paper. This adaptation improves the ability of the model to be used for prompt tuning. Note : A popular fine tuned version of the T5 Version 1.1 LM Adapted model is BigScience's T0pp. Pretraining Dataset: C4 Other Community Checkpoints: here Paper: Exploring the Limits of Transfer Learning with a Unified Text to Text Transformer Authors: Colin Raffel, Noam Shazeer, Adam Roberts, Kath…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy