🔎 KURE v1 Introducing Korea University Retrieval Embedding model, KURE v1 It has shown remarkable performance in Korean text retrieval, speficially overwhelming most multilingual embedding models. To our knowledge, It is one of the best publicly opened Korean retrieval models. For details, visit the KURE repository Model Versions Model Name Dimension Sequence Length Introduction : : : : : : : : KURE v1 1024 8192 Fine tuned BAAI/bge m3 with Korean data via CachedGISTEmbedLoss KoE5 1024 512 Fine tuned intfloat/multilingual e5 large with ko triplet v1.0 via CachedMultipleNegativesRankingLoss Model Description This is the model card of a 🤗 transformers model that has been pushed on the Hub. Developed by: NLP&AI Lab Language(s) (NLP): Korean, English License: MIT Finetuned from model: BAAI/bge m3 Example code Install Dependencies First install the Sentence Transformers library: Python code Then you can load this model and run inference. Training Details Training Data KURE v1 Korean query document hard negative(5) data 2,000,000 examples Training Procedure loss: Used CachedGISTEmbedLoss by sentence transformers batch size: 4096 learning rate: 2e 05 epochs: 1 Evaluation Metrics Recall,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy