TinyBERT: Distilling BERT for Natural Language Understanding ======== TinyBERT is 7.5x smaller and 9.4x faster on inference than BERT base and achieves competitive performances in the tasks of natural language understanding. It performs a novel transformer distillation at both the pre training and task specific learning stages. In general distillation, we use the original BERT base without fine tuning as the teacher and a large scale text corpus as the learning data. By performing the Transformer distillation on the text from general domain, we obtain a general TinyBERT which provides a good initialization for the task specific distillation. We here provide the general TinyBERT for your tasks at hand. For more details about the techniques of TinyBERT, refer to our paper: TinyBERT: Distilling BERT for Natural Language Understanding Citation ======== If you find TinyBERT useful in your research, please cite the following paper:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy