AraBERT v1 & v2 : Pre training BERT for Arabic Language Understanding AraBERT is an Arabic pretrained language model based on Google's BERT architechture. AraBERT uses the same BERT Base config. More details are available in the AraBERT Paper and in the AraBERT Meetup There are two versions of the model, AraBERTv0.1 and AraBERTv1, with the difference being that AraBERTv1 uses pre segmented text where prefixes and suffixes were split using the Farasa Segmenter. We evaluate AraBERT models on different downstream tasks and compare them to mBERT), and other state of the art models ( To the extent of our knowledge ). The Tasks were Sentiment Analysis on 6 different datasets (HARD, ASTD Balanced, ArsenTD Lev, LABR), Named Entity Recognition with the ANERcorp, and Arabic Question Answering on Arabic SQuAD and ARCD AraBERTv2 What's New! AraBERT now comes in 4 new variants to replace the old v1 versions: More Detail in the AraBERT folder and in the README and in the AraBERT Paper Model HuggingFace Model Name Size (MB/Params) Pre Segmentation DataSet (Sentences/Size/nWords) : : : : : : : : AraBERTv0.2 base bert base arabertv02 543MB / 136M No 200M / 77GB / 8.6B AraBERTv0.2 large bert large a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy