all-processeddataset is a concatenation of ofmedical-meadow-*andchatdoctor_healthcaremagicdatasets- The
ChatDoctorterm is replaced by thechatbotterm in thechatdoctor_healthcaremagicdataset - Similar to the literature the
medical_meadow_cord19dataset is subsampled to 50,000 samples truthful-qa-*is a benchmark dataset for evaluating the truthfulness of models in text generation, which is used in Llama 2 paper. Within this dataset, there are 55 and 16 questions related toHealthandNutrition, respectively, making it a valuable resource for medical question-answering scenarios.