InternVL Chat V1 2 [\[๐ GitHub\]](https://github.com/OpenGVLab/InternVL) [\[๐ InternVL 1.0\]](https://huggingface.co/papers/2312.14238) [\[๐ InternVL 1.5\]](https://huggingface.co/papers/2404.16821) [\[๐ Mini InternVL\]](https://arxiv.org/abs/2410.16261) [\[๐ InternVL 2.5\]](https://huggingface.co/papers/2412.05271) [\[๐ Blog\]](https://internvl.github.io/blog/) [\[๐จ๏ธ Chat Demo\]](https://internvl.opengvlab.com/) [\[๐ค HF Demo\]](https://huggingface.co/spaces/OpenGVLab/InternVL) [\[๐ Quick Start\]]( quick start) [\[๐ Documents\]](https://internvl.readthedocs.io/en/latest/) Introduction We are excited to introduce ๐ค InternVL Chat V1 2. Inspired by LLaVA NeXT 34B, we have also adopted Nous Hermes 2 Yi 34B as the language model. Below is the pipeline. From the experimental results, we've observed that a stronger language model (34B) can better leverage the powerful capabilities of our vision foundation model. For better training reproducibility, we follow the minimalist design and data efficiency similar to LLaVA NeXT. To reduce training costs, we provide a pre trained MLP projector and only employ around 1.2 million visual instruction tuning samples for SFT. Our model has aโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy