Llama 3.1 Nemotron Nano VL 8B V1 Model Overview Description Llama Nemotron Nano VL is a leading document intelligence vision language model (VLMs) that enables the ability to query and summarize images from the physical or virtual world. Llama Nemotron Nano VL is deployable in the data center, cloud and at the edge, including Jetson Orin and laptop by AWQ 4bit quantization through TinyChat framework. We find: (1) image text pairs are not enough, interleaved image text is essential; (2) unfreezing LLM during interleaved image text pre training enables in context learning; (3)re blending text only instruction data is crucial to boost both VLM and text only performance. This model was trained on commercial images for all three stages of training and supports single image inference. Note: NVIDIA Nemotron Nano v2 12B VL is now available on Huggingface in the BF16, FP8 and NVFP4 QAD formats. License/Terms of Use Governing Terms: Your use of the model is governed by the NVIDIA Open License Agreement. Additional Information: Llama 3.1 Community Model License; Built with Llama. Additional Information: Llama 3.1 Community Model License; Built with Llama. Deployment Geography: Global Use Case…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy