LLaVa Next, leveraging mistralai/Mistral 7B Instruct v0.2 as LLM The LLaVA NeXT model was proposed in LLaVA NeXT: Improved reasoning, OCR, and world knowledge by Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, Yong Jae Lee. LLaVa NeXT (also called LLaVa 1.6) improves upon LLaVa 1.5 by increasing the input image resolution and training on an improved visual instruction tuning dataset to improve OCR and common sense reasoning. Disclaimer: The team releasing LLaVa NeXT did not write a model card for this model so this model card has been written by the Hugging Face team. Model description LLaVa combines a pre trained large language model with a pre trained vision encoder for multimodal chatbot use cases. LLaVA 1.6 improves on LLaVA 1.5 BY: Using Mistral 7B (for this checkpoint) and Nous Hermes 2 Yi 34B which has better commercial licenses, and bilingual support More diverse and high quality data mixture Dynamic high resolution Intended uses & limitations You can use the raw model for tasks like image captioning, visual question answering, multimodal chatbot use cases. See the model hub to look for other versions on a task that interests you. How to use Here's th…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy