SauerkrautLM EuroColBERT This model is a powerful multilingual Late Interaction retriever that leverages: Continuous Pretraining with 5.4 billion English tokens using knowledge distillation from state of the art reranker models. EuroBERT Foundation building upon the multilingual EuroBERT 210m model specifically designed for European languages. Late Interaction Architecture initialized with PyLate for precise token level matching and superior retrieval performance. 🎯 Core Features and Innovations: Continuous Pretraining with Distillation : Enhanced with 5,430,249,475 English tokens while learning from powerful reranker models throughout the training process Strong Multilingual Foundation : Built on EuroBERT/EuroBERT 210m, which was specifically trained for European language understanding Full 210M Parameters : Preserving the complete capacity of the base model for maximum multilingual performance 💪 Standing on the Shoulders of Giants Starting from the exceptional EuroBERT 210m foundation – a model specifically designed for European languages – we've enhanced it with: 5.4 billion additional English tokens through continuous pretraining Knowledge distillation from state of the art r…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy