Model Card for DistilRoBERTa base Table of Contents 1. Model Details 2. Uses 3. Bias, Risks, and Limitations 4. Training Details 5. Evaluation 6. Environmental Impact 7. Citation 8. How To Get Started With the Model Model Details Model Description This model is a distilled version of the RoBERTa base model. It follows the same training procedure as DistilBERT. The code for the distillation process can be found here. This model is case sensitive: it makes a difference between english and English. The model has 6 layers, 768 dimension and 12 heads, totalizing 82M parameters (compared to 125M parameters for RoBERTa base). On average DistilRoBERTa is twice as fast as Roberta base. We encourage users of this model card to check out the RoBERTa base model card to learn more about usage, limitations and potential biases. Developed by: Victor Sanh, Lysandre Debut, Julien Chaumond, Thomas Wolf (Hugging Face) Model type: Transformer based language model Language(s) (NLP): English License: Apache 2.0 Related Models: RoBERTa base model card Resources for more information: GitHub Repository Associated Paper Uses Direct Use and Downstream Use You can use the raw model for masked language modelin…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy