Model Overview Description This family of models performs vision language and text only tasks including optical character recognition, multimodal reasoning, localization, common sense reasoning, world knowledge utilization, and coding. This model is ready for non commercial use. License/Terms of Use Governing Terms: Deed Attribution NonCommercial 4.0 International Creative Commons. Additional Information: LICENSE · Qwen/Qwen2 72B Instruct at main for Qwen2 72B Instruct and The MIT License – Open Source Initiative for InternViT 6B 448px V1 2. Model Details Today (September 17th, 2024), we introduce NVLM 1.0, a family of frontier class multimodal large language models (LLMs) that achieve state of the art results on vision language tasks, rivaling the leading proprietary models (e.g., GPT 4o) and open access models (e.g., Llama 3 V 405B and InternVL 2). Remarkably, NVLM 1.0 shows improved text only performance over its LLM backbone after multimodal training. In this repo, we are open sourcing NVLM 1.0 D 72B (decoder only architecture), the decoder only model weights and code for the community. Reference(s) Paper   Inference Code (HF)   Training Code   Website Benchmark…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy