MedQA Darija MultiLingual The largest open trilingual medical Q&A dataset with directly playable speech audio for English, French, and Moroccan Darija. A research dataset for the BRAIN HEALTH initiative, designed for multilingual medical NLP, low resource speech recognition, healthcare chatbots, and clinical education tools targeting Morocco and the broader Maghreb region. Dataset is currently in scientific validation phase . After programmatic validation (Stage 1 LOF outlier detection, Stage 2 Causal DAG consistency, Stage 3 linguistic & medical NER coverage), it will be reviewed by licensed medical professionals before peer reviewed publication. At a glance Property Value Total Q&A pairs 267,627 Specialties covered 71 Languages English · French · Moroccan Darija (Arabic script) Audio modality 6 MP3 per row (Q+A × 3 languages) Real world data 163,488 pairs (open medical corpus + iCliniq scraped) Synthetic data 104,139 pairs (multi LLM generated, validated) Audio files ~427k MP3 (~91 GB) on Hugging Face Xet storage Total file size ~94 GB Configurations default — Synthetic (104,139 rows) Trilingual synthetic Q&A generated by multiple LLM providers (Mistral, Cerebras, Groq, GPT 4o mi…
Runs entirely in your browser via DuckDB-Wasm — this dataset's real data file is loaded once, then queried locally. Nothing is sent to a server.
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy