Phi 4 reasoning plus Model Card Phi 4 reasoning Technical Report Model Summary Developers Microsoft Research Description Phi 4 reasoning plus is a state of the art open weight reasoning model finetuned from Phi 4 using supervised fine tuning on a dataset of chain of thought traces and reinforcement learning. The supervised fine tuning dataset includes a blend of synthetic prompts and high quality filtered data from public domain websites, focused on math, science, and coding skills as well as alignment data for safety and Responsible AI. The goal of this approach was to ensure that small capable models were trained with data focused on high quality and advanced reasoning. Phi 4 reasoning plus has been trained additionally with Reinforcement Learning, hence, it has higher accuracy but generates on average 50% more tokens, thus having higher latency. Architecture Base model same as previously released Phi 4, 14B parameters, dense decoder only Transformer model Inputs Text, best suited for prompts in the chat format Context length 32k tokens GPUs 32 H100 80G Training time 2.5 days Training data 16B tokens, ~8.3B unique tokens Outputs Generated text in response to the input. Model resp…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy