News [2025/04/22] We split the data and kept only the medical SFT dataset ( medical o1 sft.json ). The file medical o1 sft mix.json contains a mix of medical and general instruction data. [2025/02/22] We released the distilled dataset from Deepseek R1 based on medical verifiable problems. You can use it to initialize your models with the reasoning chain from Deepseek R1 . [2024/12/25] We open sourced the medical reasoning dataset for SFT, built on medical verifiable problems and an LLM verifier. Introduction This dataset is used to fine tune HuatuoGPT o1, a medical LLM designed for advanced medical reasoning. This dataset is constructed using GPT 4o, which searches for solutions to verifiable medical problems and validates them through a medical verifier. For details, see our paper and GitHub repository. Citation If you find our data useful, please consider citing our work!
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy