Phi 4 Reasoning Vision 15B Official Microsoft Blog Technical Report Github Try Phi 4 Reasoning Vision 15B on Microsoft Foundry Developer: Microsoft Corporation Authorized Representative: Microsoft Ireland Operations Limited, 70 Sir John Rogerson's Quay, Dublin 2, D02 R296, Ireland Release Date: March 4, 2026 License: MIT Parameters: 15B Context Length: 16,384 tokens Inputs: Text and Images Outputs: Text Training GPUs: 240 B200s Training Time: 4 days Training Dates: February 3, 2025 – February 21, 2026 Model Dependencies: Phi 4 Reasoning 1. Model Overview Phi 4 Reasoning Vision 15B is a compact open weight multimodal reasoning model built on the Phi 4 Reasoning language model backbone and the SigLIP 2 vision encoder, using a mid fusion architecture. In this architecture, the vision encoder first converts images into visual tokens, which are then projected into the language model's embedding space and injected into the pretrained language model. This approach leverages the strengths of both pretrained components while keeping training and inference costs manageable. The model employs a dynamic resolution vision encoder with up to 3,600 visual tokens, enabling high resolution image un…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy