Mursit Base TR Retrieval Model Description Mursit Base TR Retrieval is a Turkish embedding model pre trained entirely from scratch on Turkish dominant corpora and fine tuned for retrieval tasks. The model is based on ModernBERT base architecture (155M parameters) and optimized specifically for Turkish legal domain applications. This model demonstrates that trainable Masked Language Modeling (MLM) models can effectively serve as foundations for embedding tasks when training quality is assessed through downstream performance rather than MLM loss minimization alone. Key Features: Pre trained from scratch on approximately 112.7 billion tokens of Turkish dominant corpus Post trained for embedding tasks using contrastive learning on MS MARCO TR dataset Achieves strong performance on Turkish legal retrieval benchmarks (55.86 MTEB Score, 47.52 Legal Score) Optimized for Turkish legal domain with custom tokenizer trained on legal documents Model Type: Embedding Parameters: 155M Base Model: newmindai/Mursit Base Architecture: ModernBERT base Embedding Dimension: 768 Max Sequence Length: 1,024 tokens Architecture Details The model is based on ModernBERT architecture, which incorporates modern…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy