EmbeddingGemma 300m finetuned on the Medical Instruction and RetrIeval Dataset (MIRIAD) This is a sentence transformers model finetuned from google/embeddinggemma 300m on the miriad/miriad 4.4M dataset (specifically the first 100.000 question passage pairs from tomaarsen/miriad 4.4M split). It maps sentences & documents to a 768 dimensional dense vector space and can be used for medical information retrieval, specifically designed for searching for passages (up to 1k tokens) of scientific medical papers using detailed medical questions. This model has been trained using code from our EmbeddingGemma blogpost to showcase how the EmbeddingGemma model can be finetuned on specific domains/tasks for even stronger performance. It is not affiliated with Google. Model Details Model Description Model Type: Sentence Transformer Base model: google/embeddinggemma 300m Maximum Sequence Length: 1024 tokens Output Dimensionality: 768 dimensions Similarity Function: Cosine Similarity Training Dataset: miriad 4.4 m split (the first 100.000 samples of the default subset) Language: en License: apache 2.0 Model Sources Documentation: Sentence Transformers Documentation Repository: Sentence Transformers…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy