deberta v3 xsmall zeroshot v1.1 all 33 This model was fine tuned using the same pipeline as described in the model card for MoritzLaurer/deberta v3 large zeroshot v1.1 all 33 and in this paper. The foundation model is microsoft/deberta v3 xsmall. The model only has 22 million backbone parameters and 128 million vocabulary parameters. The backbone parameters are the main parameters active during inference, providing a significant speedup over larger models. The model is 142 MB small. This model was trained to provide a small and highly efficient zeroshot option, especially for edge devices or in browser use cases with transformers.js. Usage and other details For usage instructions and other details refer to this model card MoritzLaurer/deberta v3 large zeroshot v1.1 all 33 and this paper. Metrics: I didn't not do zeroshot evaluation for this model to save time and compute. The table below shows standard accuracy for all datasets the model was trained on (note that the NLI datasets are binary). General takeaway: the model is much more efficient than its larger sisters, but it performs less well. Datasets mnli m mnli mm fevernli anli r1 anli r2 anli r3 wanli lingnli wellformedquery ro…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy