Releasing zeroentropy/zembed 1 In retrieval systems, embedding models determine the quality of your search. However, SOTA embedding models are closed source and proprietary. At ZeroEntropy, we've trained a SOTA 4B open weight multilingual embedding model that outperforms every competitor we benchmarked, and we're launching it here on HuggingFace. This model outperforms OpenAI text embedding large , Cohere Embed v4 , gemini embedding 001 , and voyage 4 nano across finance, healthcare, legal, conversational, manufacturing, code, and STEM. zembed 1 is distilled directly from our SOTA reranker zerank 2 using our zELO methodology, which models relevance scores as adjusted Elo ratings. Standard contrastive training on binary labels can't match this signal. See our blog post for details. The model supports flexible dimension projections (2560, 1280, 640, 320, 160, 80, 40) and quantization down to binary, compressing a full 8 KB vector to under 128 bytes with a controlled accuracy trade off. See our Technical Report (Coming soon!) for details on the projection method. zembed 1 is multilingual from the ground up, with over half the training data in non English languages. This model is relea…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy