Arabic Triplet Matryoshka V2 Model [ATM2] Model Description Arabic Triplet Matryoshka V2 Model is a state of the art Arabic language embedding model based on the sentence transformers framework. It is fine tuned from aubmindlab/bert base arabertv02 and specifically designed to capture the rich semantic nuances of Arabic text. It is described in detail in the paper GATE: General Arabic Text Embedding for Enhanced Semantic Textual Similarity with Hybrid Loss Training. This model maps sentences and paragraphs to a 768 dimensional dense vector space, enabling high quality semantic text operations including: Semantic textual similarity Semantic search Paraphrase mining Text classification Clustering Information retrieval Question answering Key Features State of the Art Performance : Achieved 0.85 on STS17 and 0.64 on STS22.v2 with an average score of 74.5, making it the leading Arabic embedding model currently available. MatryoshkaLoss Training : Utilizes nested embedding learning techniques to create hierarchical embeddings at multiple resolutions. Optimization : Trained for 3 epochs with a final training loss of 0.718. Full Arabic Language Support : Designed specifically to handle the…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy