OpenVLA 7B OpenVLA 7B ( openvla 7b ) is an open vision language action model trained on 970K robot manipulation episodes from the Open X Embodiment dataset. The model takes language instructions and camera images as input and generates robot actions. It supports controlling multiple robots out of the box, and can be quickly adapted for new robot domains via (parameter efficient) fine tuning. All OpenVLA checkpoints, as well as our training codebase are released under an MIT License. For full details, please read our paper and see our project page. Model Summary Developed by: The OpenVLA team consisting of researchers from Stanford, UC Berkeley, Google Deepmind, and the Toyota Research Institute. Model type: Vision language action (language, image = robot actions) Language(s) (NLP): en License: MIT Finetuned from: prism dinosiglip 224px , a VLM trained from: + Vision Backbone : DINOv2 ViT L/14 and SigLIP ViT So400M/14 + Language Model : Llama 2 Pretraining Dataset: Open X Embodiment specific component datasets can be found here. Repository: https://github.com/openvla/openvla Paper: OpenVLA: An Open Source Vision Language Action Model Project Page & Videos: https://openvla.github.io/…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy