Model Card for DistilBERT base multilingual (cased) Table of Contents 1. Model Details 2. Uses 3. Bias, Risks, and Limitations 4. Training Details 5. Evaluation 6. Environmental Impact 7. Citation 8. How To Get Started With the Model Model Details Model Description This model is a distilled version of the BERT base multilingual model. The code for the distillation process can be found here. This model is cased: it does make a difference between english and English. The model is trained on the concatenation of Wikipedia in 104 different languages listed here. The model has 6 layers, 768 dimension and 12 heads, totalizing 134M parameters (compared to 177M parameters for mBERT base). On average, this model, referred to as DistilmBERT, is twice as fast as mBERT base. We encourage potential users of this model to check out the BERT base multilingual model card to learn more about usage, limitations and potential biases. Developed by: Victor Sanh, Lysandre Debut, Julien Chaumond, Thomas Wolf (Hugging Face) Model type: Transformer based language model Language(s) (NLP): 104 languages; see full list here License: Apache 2.0 Related Models: BERT base multilingual model Resources for more in…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy