OpenVLA OFT LIBERO (GGUF for vla.cpp) GGUF conversion of moojink/openvla 7b oft finetuned libero spatial object goal 10 for inference with vla.cpp , a lightweight C++ inference engine for Vision Language Action models built on top of llama.cpp . OpenVLA OFT (Optimized Fine Tuning) is a 7B scale VLA: a Prismatic style VLM with a fused DINOv2 L/14 reg4 + SigLIP SO400m/14 vision backbone (224 px) over a Llama 2 7B language backbone, coupled to an MLPResNet (L1 regression) action head that predicts an 8 step action chunk in parallel (no autoregressive decoding). It consumes two camera views (third person + wrist) and an 8 D proprioceptive state. This checkpoint is a single multi task fine tune covering all four LIBERO suites spatial, object, goal, and 10 (LIBERO Long) . The vision tower is baked into the combined GGUF, so no separate mmproj file is needed . Files File Size Description : openvla oft libero.gguf 14.03 GiB Combined VLA model fused DINOv2+SigLIP vision tower + Llama 2 7B LM + MLPResNet L1 action head + proprio projector + dataset stats + arch config, BF16 dataset statistics.json Action/state normalisation stats for all four suites (required by the client) Usage Build vla s…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy