Kvasir VQA x1 A Multimodal Dataset for Medical Reasoning and Robust MedVQA in Gastrointestinal Endoscopy Kvasir VQA x1 on GitHub Original Image from Kvasir VQA(Simula Datasets) Paper 🔗 MediaEval Medico 2025 Challenge uses this dataset. We encourage you to check out and participate! Overview Kvasir VQA x1 is a large scale dataset designed to benchmark medical visual question answering (MedVQA) in gastrointestinal (GI) endoscopy. It introduces 159,549 new QA pairs stratified by clinical complexity, along with support for visual robustness testing via augmentations. Features Each dataset entry includes: img id : Unique reference to an image from Kvasir VQA complexity : Question complexity level (1–3) question : Complex, natural language question answer : Clinically validated answer original : List of atomic QA pairs merged into the complex question question class : Associated clinical category labels Splits train : Training samples only test : Held out samples for final evaluation To ensure generalization, no image or QA from the test set appears in the training set. Downloading Images and Preparing JSONLs 🔗 Instructions and code for creating VLM ready JSONLs that incorporates the a…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy