InternVL3 2B Transformers π€ Implementation [\[π InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[π InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[π InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[π InternVL2.5 MPO\]](https://huggingface.co/papers/2411.10442) [\[π InternVL3\]](https://huggingface.co/papers/2504.10479) [\[π Blog\]](https://internvl.github.io/blog/) [\[π¨οΈ Chat Demo\]](https://internvl.opengvlab.com/) [\[π€ HF Demo\]](https://huggingface.co/spaces/OpenGVLab/InternVL) [\[π Quick Start\]]( quick start) [\[π Documents\]](https://internvl.readthedocs.io/en/latest/) [!IMPORTANT] This repository contains the Hugging Face π€ Transformers implementation for the OpenGVLab/InternVL3 2B model. It is intended to be functionally equivalent to the original OpenGVLab release. As a native Transformers model, it supports core library features such as various attention implementations (eager, including SDPA, and FA2) and enables efficient batched inference with interleaved image, video, and text inputs. Introduction We introduce InternVL3, an advanced multimodal large language model (MLLM) series that demonstrates superior overall performβ¦
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy