PubMedBERT Embeddings This is a PubMedBERT base model fined tuned using sentence transformers. It maps sentences & paragraphs to a 768 dimensional dense vector space and can be used for tasks like clustering or semantic search. The training dataset was generated using a random sample of PubMed title abstract pairs along with similar title pairs. PubMedBERT Embeddings produces higher quality embeddings than generalized models for medical literature. Further fine tuning for a medical subdomain will result in even better performance. Usage (txtai) This model can be used to build embeddings databases with txtai for semantic search and/or as a knowledge source for retrieval augmented generation (RAG). Usage (Sentence Transformers) Alternatively, the model can be loaded with sentence transformers. Usage (Hugging Face Transformers) The model can also be used directly with Transformers. Evaluation Results Performance of this model compared to the top base models on the MTEB leaderboard is shown below. A popular smaller model was also evaluated along with the most downloaded PubMed similarity model on the Hugging Face Hub. The following datasets were used to evaluate model performance. PubM…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy