DistilBERT base cased distilled SQuAD Table of Contents Model Details How To Get Started With the Model Uses Risks, Limitations and Biases Training Evaluation Environmental Impact Technical Specifications Citation Information Model Card Authors Model Details Model Description: The DistilBERT model was proposed in the blog post Smaller, faster, cheaper, lighter: Introducing DistilBERT, adistilled version of BERT, and the paper DistilBERT, adistilled version of BERT: smaller, faster, cheaper and lighter. DistilBERT is a small, fast, cheap and light Transformer model trained by distilling BERT base. It has 40% less parameters than bert base uncased , runs 60% faster while preserving over 95% of BERT's performances as measured on the GLUE language understanding benchmark. This model is a fine tune checkpoint of DistilBERT base cased, fine tuned using (a second step of) knowledge distillation on SQuAD v1.1. Developed by: Hugging Face Model Type: Transformer based language model Language(s): English License: Apache 2.0 Related Models: DistilBERT base cased Resources for more information: See this repository for more about Distil\ (a class of compressed models including this model) See Sa…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy