Qwen3.5 4B Medical GSPO A Chinese medical reasoning model fine tuned from Qwen3.5 4B using a two stage training pipeline: Supervised Fine Tuning (SFT) for format alignment, followed by Group Sequence Policy Optimization (GSPO) with an LLM as Judge reward function. Model Description This model is designed to produce structured chain of thought (CoT) reasoning for Chinese medical questions, including clinical diagnosis, treatment planning, and differential diagnosis. Training Pipeline Stage 1 — Supervised Fine Tuning (SFT) The model was first trained with SFT on the FreedomIntelligence/medical o1 reasoning SFT dataset (Chinese subset) to establish a consistent output format: a ... reasoning block followed by a concise final answer. Stage 2 — GSPO with LLM as Judge GSPO was proposed by Zheng et al. (arXiv:2507.18071). Reinforcement learning was applied using GSPO (Group Sequence Policy Optimization), a sequence level variant of GRPO that computes importance ratios and clipping at the sequence level rather than the token level, improving training stability over long horizons. The reward function uses DeepSeek Chat as an LLM judge with a 5 tier scoring scheme: Score Criterion +2.0 Same…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy