Model Card for Video LLaVa Model Details Model type: Video LLaVA is an open source multomodal model trained by fine tuning LLM on multimodal instruction following data. It is an auto regressive language model, based on the transformer architecture. Base LLM: lmsys/vicuna 13b v1.5 Model Description: The model can generate interleaving images and videos, despite the absence of image video pairs in the dataset. Video LLaVa is uses an encoder trained for unified visual representation through alignment prior to projection. Extensive experiments demonstrate the complementarity of modalities, showcasing significant superiority when compared to models specifically designed for either images or videos. VideoLLaVa example. Taken from the original paper. Paper or resources for more information: https://github.com/PKU YuanGroup/Video LLaVA 🗝️ Training Dataset The images pretraining dataset is from LLaVA. The images tuning dataset is from LLaVA. The videos pretraining dataset is from Valley. The videos tuning dataset is from Video ChatGPT. How to Get Started with the Model Use the code below to get started with the model. 👍 Acknowledgement LLaVA The codebase we built upon and it is an efficie…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy