LLaVA Video 7B Qwen2 Table of Contents 1. Model Summary 2. Use 3. Limitations 4. Training 5. License 6. Citation Model Summary The LLaVA Video models are 7/72B parameter models trained on LLaVA Video 178K and LLaVA OneVision Dataset, based on Qwen2 language model with a context window of 32K tokens. This model support at most 64 frames. Project Page: Project Page. Paper : For more details, please check our paper Repository: LLaVA VL/LLaVA NeXT Point of Contact: Yuanhan Zhang Languages: English, Chinese Use Intended use The model was trained on LLaVA Video 178K and LLaVA OneVision Dataset, having the ability to interact with images, multi image and videos, but specific to videos. Feel free to share your generations in the Community tab! Generation We provide the simple generation process for using our model. For more details, you could refer to Github. Training Model Architecture: SO400M + Qwen2 Initialized Model: lmms lab/llava onevision qwen2 7b si Data: A mixture of 1.6M single image/multi image/video data, 1 epoch, full model Precision: bfloat16 Hardware & Software GPUs: 256 Nvidia Tesla A100 (for whole model series training) Orchestration: Huggingface Trainer Neural networks: P…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy