nanoLLaVA Sub 1B Vision Language Model Description nanoLLaVA is a "small but mighty" 1B vision language model designed to run efficiently on edge devices. Base LLM : Quyen SE v0.1 (Qwen1.5 0.5B) Vision Encoder : google/siglip so400m patch14 384 Model VQA v2 TextVQA ScienceQA POPE MMMU (Test) MMMU (Eval) GQA MM VET Score 70.84 46.71 58.97 84.1 28.6 30.4 54.79 23.9 Training Data Training Data will be released later as I am still writing a paper on this. Expect the final final to be much more powerful than the current one. Finetuning Code Coming Soon!!! Usage You can use with transformers with the following script: Prompt Format The model follow the ChatML standard, however, without \n at the end of : Image Example What is the text saying? "Small but mighty". How does the text correlate to the context of the image? The text seems to be a playful or humorous representation of a small but mighty figure, possibly a mouse or a mouse toy, holding a weightlifting bar. Model is trained using a modified version from Bunny
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy