ember v1 This model has been trained on an extensive corpus of text pairs that encompass a broad spectrum of domains, including finance, science, medicine, law, and various others. During the training process, we incorporated techniques derived from the RetroMAE and SetFit research papers. Plans The research paper will be published soon. The v2 of the model is currently in development and will feature an extended maximum sequence length of 4,000 tokens. Usage Use with transformers: Use with sentence transformers: Massive Text Embedding Benchmark (MTEB) Evaluation Our model achieve state of the art performance on MTEB leaderboard Model Name Dimension Sequence Length Average (56) : : : : : : : : ember v1 1024 512 63.54 bge large en v1.5 1024 512 63.23 bge base en v1.5 768 512 63.05 text embedding ada 002 1536 8191 60.99 Limitation This model exclusively caters to English texts, and any lengthy texts will be truncated to a maximum of 512 tokens. License MIT Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy