Dream VL 7B Dream VL 7B is an open diffusion vision language model trained on 12M multimodal data from the MAmmoTH VL Instruct 12M dataset. The model takes language instructions and images as input and generates language outputs. All Dream VL checkpoints, as well as our training codebase are released under an Apache 2.0 License. For full details, please read our blog and the paper: Dream VL & Dream VLA: Open Vision Language and Vision Language Action Models with Diffusion Language Model Backbone. Model Summary Model type: Vision language (language, image = language) Language(s) (NLP): en License: apache 2.0 Finetuned from: Dream 7B , with Qwen2ViT Vision Backbone. Pretraining Dataset: MAmmoTH VL Instruct 12M. Repository: https://github.com/DreamLM/Dream VLX Project Page & Videos: https://hkunlp.github.io/blog/2025/dream vlx Getting Started Citation BibTeX:
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy