PubMedBERT SPLADE This is a SPLADE Sparse Encoder model finetuned from PubMedBERT base using sentence transformers. It maps sentences & paragraphs to a 30522 dimensional sparse vector space and can be used for semantic search and sparse retrieval. The training dataset was generated using a random sample of PubMed title abstract pairs along with similar title pairs. PubMedBERT SPLADE produces higher quality sparse embeddings than generalized models for medical literature. Further fine tuning for a medical subdomain will result in even better performance. Usage (txtai) This model can be used to build embeddings databases with txtai for semantic search and/or as a knowledge source for retrieval augmented generation (RAG). Note: txtai 9.0+ is required for sparse vector scoring support Usage (Sentence Transformers) Alternatively, the model can be loaded with sentence transformers. Evaluation Results Performance of this model compared to the top base models on the MTEB leaderboard is shown below. A popular smaller model was also evaluated along with the most downloaded PubMed similarity model on the Hugging Face Hub. The following datasets were used to evaluate model performance. PubMed…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy