GATE AraBert V1 This is GATE General Arabic Text Embedding trained using SentenceTransformers in a multi task setup. The system trains on the AllNLI and on the STS dataset. It is described in detail in the paper GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Hybrid Loss Training. Project page: https://huggingface.co/collections/Omartificial Intelligence Space/arabic matryoshka embedding models 666f764d3b570f44d7f77d4e Model Details Model Description Model Type: Sentence Transformer Base model: Omartificial Intelligence Space/Arabic Triplet Matryoshka V2 Maximum Sequence Length: 512 tokens Output Dimensionality: 768 tokens Similarity Function: Cosine Similarity Training Datasets: all nli sts Language: ar Usage Direct Usage (Sentence Transformers) First install the Sentence Transformers library: Then you can load this model and run inference. Evaluation Model Dim Params. STS17 STS22 v2 Average Arabic Triplet Matryoshka V2 768 135M 85 64 75 Arabert all nli triplet Matryoshka 768 135M 83 64 74 AraGemma Embedding 300m 768 303M 84 62 73 GATE AraBert V1 767 135M 83 63 73 Marbert all nli triplet Matryoshka 768 163M 82 61 72 Arabic labse Matryoshka 768 471…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy