LLaVa Next Model Card The LLaVA NeXT model was proposed in LLaVA NeXT: Stronger LLMs Supercharge Multimodal Capabilities in the Wild by Bo Li, Kaichen Zhang, Hao Zhang, Dong Guo, Renrui Zhang, Feng Li, Yuanhan Zhang, Ziwei Liu, Chunyuan Li. These LLaVa NeXT series improves upon LLaVa 1.6 by training with stringer language backbones, improving the performance. Disclaimer: The team releasing LLaVa NeXT did not write a model card for this model so this model card has been written by the Hugging Face team. Model description LLaVa combines a pre trained large language model with a pre trained vision encoder for multimodal chatbot use cases. LLaVA NeXT Llama3 improves on LLaVA 1.6 BY: More diverse and high quality data mixture Better and bigger language backbone Base LLM: meta llama/Meta Llama 3 8B Instruct Intended uses & limitations You can use the raw model for tasks like image captioning, visual question answering, multimodal chatbot use cases. See the model hub to look for other versions on a task that interests you. How to use To run the model with the pipeline , see the below example: You can also load and use the model like following: From transformers =v4.48, you can also pass i…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy