LLaVA OneVision 1.5: Fully Open Source State of the Art VLM Model Introduction LLaVA OneVision 1.5 is a fully open source family of large multimodal models (LMMs) built to democratize multimodal training. Trained on native‑resolution images, it delivers state‑of‑the‑art performance at substantially lower cost. The project also releases high‑quality pretraining and SFT data, a complete and efficient training framework with recipes and configs, and comprehensive logs to support transparent, reproducible research. Superior Performance The model leads on multiple multimodal benchmarks and generally surpasses Qwen2.5 VL. Training on native resolution images significantly improves its visual understanding. High Quality Data at Scale The pretraining corpus comprises large scale, concept balanced, diverse, and high quality captions curated with strict filtering and quality control. The instruction tuning dataset is comprehensive and covers a wide range of tasks. Ultra Efficient Training Framework The end to end training cost is about $16,000 on A100 GPUs at roughly $0.60 per GPU hour. The system is built on Megatron LM with support for MoE, FP8, and long sequence parallelism, and the codeb…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy