π Phi 4 : [mini reasoning reasoning] [multimodal instruct onnx]; [mini instruct onnx] Model Summary Phi 4 multimodal instruct is a lightweight open multimodal foundation model that leverages the language, vision, and speech research and datasets used for Phi 3.5 and 4.0 models. The model processes text, image, and audio inputs, generating text outputs, and comes with 128K token context length. The model underwent an enhancement process, incorporating both supervised fine tuning, direct preference optimization and RLHF (Reinforcement Learning from Human Feedback) to support precise instruction adherence and safety measures. The languages that each modal supports are the following: Text: Arabic, Chinese, Czech, Danish, Dutch, English, Finnish, French, German, Hebrew, Hungarian, Italian, Japanese, Korean, Norwegian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, Turkish, Ukrainian Vision: English Audio: English, Chinese, German, French, Italian, Japanese, Spanish, Portuguese π° Phi 4 multimodal Microsoft Blog π Phi 4 multimodal Technical Report π‘ Phi Portal π©βπ³ Phi Cookbook π₯οΈ Try It on Azure, GitHub, Nvidia, Huggingface playgrounds π±Huggingface Spaces Thoughts Organizer,β¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy