Model card for mDeBERTa v3 base xnli multilingual nli 2mil7 Model description This multilingual model can perform natural language inference (NLI) on 100 languages and is therefore also suitable for multilingual zero shot classification. The underlying mDeBERTa v3 base model was pre trained by Microsoft on the CC100 multilingual dataset with 100 languages. The model was then fine tuned on the XNLI dataset and on the multilingual NLI 26lang 2mil7 dataset. Both datasets contain more than 2.7 million hypothesis premise pairs in 27 languages spoken by more than 4 billion people. As of December 2021, mDeBERTa v3 base is the best performing multilingual base sized transformer model introduced by Microsoft in this paper. How to use the model Simple zero shot classification pipeline NLI use case Training data This model was trained on the multilingual nli 26lang 2mil7 dataset and the XNLI validation dataset. The multilingual nli 26lang 2mil7 dataset contains 2 730 000 NLI hypothesis premise pairs in 26 languages spoken by more than 4 billion people. The dataset contains 105 000 text pairs per language. It is based on the English datasets MultiNLI, Fever NLI, ANLI, LingNLI and WANLI and was…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy