The PubMed Corpus in MedRAG This HF dataset contains the snippets from the PubMed corpus used in MedRAG. It can be used for medical Retrieval Augmented Generation (RAG). News (02/26/2024) The "id" column has been reformatted. A new "PMID" column is added. Dataset Details Dataset Descriptions PubMed is the most widely used literature resource, containing over 36 million biomedical articles. For MedRAG, we use a PubMed subset of 23.9 million articles with valid titles and abstracts. This HF dataset contains our ready to use snippets for the PubMed corpus, including 23,898,701 snippets with an average of 296 tokens. Dataset Structure Each row is a snippet of PubMed, which includes the following features: id: a unique identifier of the snippet title: the title of the PubMed article from which the snippet is collected content: the abstract of the PubMed article from which the snippet is collected contents: a concatenation of 'title' and 'content', which will be used by the BM25 retriever Uses Direct Use Use in MedRAG Citation
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy