Model Summary Phi 3.5 vision is a lightweight, state of the art open multimodal model built upon datasets which include synthetic data and filtered publicly available websites with a focus on very high quality, reasoning dense data both on text and vision. The model belongs to the Phi 3 model family, and the multimodal version comes with 128K context length (in tokens) it can support. The model underwent a rigorous enhancement process, incorporating both supervised fine tuning and direct preference optimization to ensure precise instruction adherence and robust safety measures. 🏡 Phi 3 Portal 📰 Phi 3 Microsoft Blog 📖 Phi 3 Technical Report 👩🍳 Phi 3 Cookbook 🖥️ Try It Phi 3.5 : [[mini instruct]](https://huggingface.co/microsoft/Phi 3.5 mini instruct); [[MoE instruct]](https://huggingface.co/microsoft/Phi 3.5 MoE instruct) ; [[vision instruct]](https://huggingface.co/microsoft/Phi 3.5 vision instruct) Intended Uses Primary Use Cases The model is intended for broad commercial and research use in English. The model provides uses for general purpose AI systems and applications with visual and text input capabilities which require: 1) Memory/compute constrained environments 2) Lat…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy