AraBERTv0.2 Twitter AraBERTv0.2 Twitter base/large are two new models for Arabic dialects and tweets, trained by continuing the pre training using the MLM task on ~60M Arabic tweets (filtered from a collection on 100M). The two new models have had emojies added to their vocabulary in addition to common words that weren't at first present. The pre training was done with a max sentence length of 64 only for 1 epoch. AraBERT is an Arabic pretrained language model based on Google's BERT architechture. AraBERT uses the same BERT Base config. More details are available in the AraBERT Paper and in the AraBERT Meetup Other Models Model HuggingFace Model Name Size (MB/Params) Pre Segmentation DataSet (Sentences/Size/nWords) : : : : : : : : AraBERTv0.2 base bert base arabertv02 543MB / 136M No 200M / 77GB / 8.6B AraBERTv0.2 large bert large arabertv02 1.38G / 371M No 200M / 77GB / 8.6B AraBERTv2 base bert base arabertv2 543MB / 136M Yes 200M / 77GB / 8.6B AraBERTv2 large bert large arabertv2 1.38G / 371M Yes 200M / 77GB / 8.6B AraBERTv0.1 base bert base arabertv01 543MB / 136M No 77M / 23GB / 2.7B AraBERTv1 base bert base arabert 543MB / 136M Yes 77M / 23GB / 2.7B AraBERTv0.2 Twitter base be…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy