CAMeLBERT: A collection of pre trained models for Arabic NLP tasks Model description CAMeLBERT is a collection of BERT models pre trained on Arabic texts with different sizes and variants. We release pre trained language models for Modern Standard Arabic (MSA), dialectal Arabic (DA), and classical Arabic (CA), in addition to a model pre trained on a mix of the three. We also provide additional models that are pre trained on a scaled down set of the MSA variant (half, quarter, eighth, and sixteenth). The details are described in the paper "The Interplay of Variant, Size, and Task Type in Arabic Pre trained Language Models." This model card describes CAMeLBERT Mix ( bert base arabic camelbert mix ), a model pre trained on a mixture of these variants: MSA, DA, and CA. Model Variant Size Word : : : : ✔ bert base arabic camelbert mix CA,DA,MSA 167GB 17.3B bert base arabic camelbert ca CA 6GB 847M bert base arabic camelbert da DA 54GB 5.8B bert base arabic camelbert msa MSA 107GB 12.6B bert base arabic camelbert msa half MSA 53GB 6.3B bert base arabic camelbert msa quarter MSA 27GB 3.1B bert base arabic camelbert msa eighth MSA 14GB 1.6B bert base arabic camelbert msa sixteenth MSA 6GB 7…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy