Qianfan VL: Domain Enhanced Universal Vision Language Models Domain Capability Enhancement through Continuous Pre training 3B to 70B Parameter Scale Document Understanding & OCR Enhancement Chain of Thought Reasoning Support This repository contains models presented in the paper Qianfan OCR: A Unified End to End Model for Document Intelligence. ๐ Quick Links Repository : ๐ป GitHub Models : ๐ค Hugging Face ๐ค ModelScope Documentation : ๐ Cookbook ๐ Technical Report Blogs : ๐จ๐ณ ไธญๆๅๅฎข ๐ฌ๐ง English Blog Model Description Qianfan VL is a series of general purpose multimodal large language models enhanced for enterprise level multimodal applications. The models offer deep optimization for high frequency scenarios in industrial deployment while maintaining strong general capabilities. Model Variants Model Parameters Context Length CoT Support Best For Qianfan VL 3B 3B 32k โ Edge deployment, real time OCR Qianfan VL 8B 8B 32k โ Server side general scenarios, fine tuning Qianfan VL 70B 70B 32k โ Complex reasoning, data synthesis Architecture Language Model : Qianfan VL 3B: Based on Qwen2.5 3B Qianfan VL 8B/70B: Based on Llama 3.1 architecture Enhanced with 3T multilingual corpus Vision Eโฆ
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy