LLaVA Onevision Model Card Check out also the Google Colab demo to run Llava on a free tier Google Colab instance: Below is the model card of 0.5B LLaVA Onevision model which is copied from the original LLaVA Onevision model card that you can find here. Model details Model type: LLaVA Onevision is an open source multimodal LLM trained by fine tuning Qwen2 on GPT generated multimodal instruction following data. LLaVA OneVision is the first single model that can simultaneously push the performance boundaries of open LMMs in three important computer vision scenarios: single image, multi image, and video scenarios. Importantly, the design of LLaVA OneVision allows strong transfer learning across different modalities/scenarios, yielding new emerging capabilities. In particular, strong video understanding and cross scenario capabilities are demonstrated through task transfer from images to videos. Model date: LLaVA Onevision 0.5 ov was added in August 2024. Paper or resources for more information: https://llava vl.github.io/ Architecture: SO400M + Qwen2 Pretraining Stage: LCS 558K, 1 epoch, projector Mid Stage: A mixture of 4.7M high quality synthetic data, 1 epoch, full model Final Imag…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy