1. Introduction Introducing DeepSeek VL, an open source Vision Language (VL) Model designed for real world vision and language understanding applications. DeepSeek VL possesses general multimodal understanding capabilities, capable of processing logical diagrams, web pages, formula recognition, scientific literature, natural images, and embodied intelligence in complex scenarios. DeepSeek VL: Towards Real World Vision Language Understanding Github Repository Haoyu Lu , Wen Liu , Bo Zhang , Bingxuan Wang, Kai Dong, Bo Liu, Jingxiang Sun, Tongzheng Ren, Zhuoshu Li, Hao Yang, Yaofeng Sun, Chengqi Deng, Hanwei Xu, Zhenda Xie, Chong Ruan ( Equal Contribution, Project Lead) 2. Model Summary DeepSeek VL 7b base uses the SigLIP L and SAM B as the hybrid vision encoder supporting 1024 x 1024 image input and is constructed based on the DeepSeek LLM 7b base which is trained on an approximate corpus of 2T text tokens. The whole DeepSeek VL 7b base model is finally trained around 400B vision language tokens. DeekSeel VL 7b chat is an instructed version based on DeepSeek VL 7b base. 3. Quick Start Installation On the basis of Python = 3.8 environment, install the necessary dependencies by runnin…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy