GIST Embedding v0 GISTEmbed: Guided In sample Selection of Training Negatives for Text Embedding Fine tuning The model is fine tuned on top of the BAAI/bge base en v1.5 using the MEDI dataset augmented with mined triplets from the MTEB Classification training dataset (excluding data from the Amazon Polarity Classification task). The model does not require any instruction for generating embeddings. This means that queries for retrieval tasks can be directly encoded without crafting instructions. Technical paper: GISTEmbed: Guided In sample Selection of Training Negatives for Text Embedding Fine tuning Data The dataset used is a compilation of the MEDI and MTEB Classification training datasets. Third party datasets may be subject to additional terms and conditions under their associated licenses. A HuggingFace Dataset version of the compiled dataset, and the specific revision used to train the model, is available: Dataset: avsolatorio/medi data mteb avs triplets Revision: 238a0499b6e6b690cc64ea56fde8461daa8341bb The dataset contains a task type key, which can be used to select only the mteb classification tasks (prefixed with mteb ). The MEDI Dataset is published in the following pap…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy