X VLA 0.9B (Foundation Edition) Repository: 2toINF/X VLA Authors: 2toINF License: Apache 2.0 Paper: Zheng et al., 2025, “X VLA: Soft Prompted Transformer as Scalable Cross Embodiment Vision Language Action Model” (arXiv:2510.10274) 🚀 Overview Successful generalist Vision Language Action (VLA) models rely on effective training across diverse robotic platforms with large scale, cross embodiment, heterogeneous datasets. To facilitate and leverage the heterogeneity in rich robotic data sources, X VLA introduces a Soft Prompt approach with minimally added parameters: we infuse prompt learning concepts into cross embodiment robot learning, introducing separate sets of learnable embeddings for each distinct embodiment. These embodiment specific prompts empower VLA models to exploit cross embodiment features effectively. Our architecture— a clean, flow matching based VLA design relying exclusively on soft prompted standard Transformers —achieves superior scalability and simplicity. Trained on Bridge Data and evaluated across six simulations and three real world robots , the 0.9B parameter X VLA simultaneously achieves state of the art performance across diverse benchmarks, demonstrating f…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy