PaliGemma model card Model page: PaliGemma Transformers PaliGemma 3B weights, fine tuned with 224 224 input images and 256 token input/output text sequences on a mixture of downstream academic datasets. The models are available in float32, bfloat16 and float16 format for research purposes only. Resources and technical documentation: Responsible Generative AI Toolkit PaliGemma on Kaggle PaliGemma on Vertex Model Garden Terms of Use: Terms Authors: Google Model information Model summary Description PaliGemma is a versatile and lightweight vision language model (VLM) inspired by PaLI 3 and based on open components such as the SigLIP vision model and the Gemma language model. It takes both image and text as input and generates text as output, supporting multiple languages. It is designed for class leading fine tune performance on a wide range of vision language tasks such as image and short video caption, visual question answering, text reading, object detection and object segmentation. Model architecture PaliGemma is the composition of a Transformer decoder and a Vision Transformer image encoder, with a total of 3 billion params. The text decoder is initialized from Gemma 2B. The imag…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy