LLaVA OneVision Play with the model on the LLaVA OneVision Chat. Table of Contents 1. Model Summary 2. Use 3. Limitations 4. Training 5. License 6. Citation Model Summary The LLaVA OneVision models are 0.5/7/72B parameter models trained on LLaVA OneVision, based on Qwen2 language model with a context window of 32K tokens. Repository: LLaVA VL/LLaVA NeXT Project Website: llava onevision.lmms lab.com Paper: LLaVA OneVision Point of Contact: Bo Li Languages: English, Chinese Use Intended use The model was trained on LLaVA OneVision Dataset and have the ability to interact with images, multi image and videos. Feel free to share your generations in the Community tab! Generation We provide the simple generation process for using our model. For more details, you could refer to Github. Training Model Architecture: SO400M + Qwen2 Pretraining Stage: LCS 558K, 1 epoch, projector Mid Stage: A mixture of 4.7M high quality synthetic data, 1 epoch, full model Final Image Stage: A mixture of 3.6M single image data, 1 epoch, full model OneVision Stage: A mixture of 1.6M single image/multi image/video data, 1 epoch, full model Precision: bfloat16 Hardware & Software GPUs: 256 Nvidia Tesla A100 (for…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy