Salience 1.5 — Flash A 30B A3B Mixture of Experts multimodal agent — only 3.3B active params per token: the decode speed of a small model with the reach of a large one. Vection Labs Weights · Benchmarks · Quickstart · Fast inference · Limitations Abstract Salience 1.5 Flash is a sparse Mixture of Experts vision language model: 30B total parameters, but only 3.3B active per token . It decodes at the speed of a ~3B dense model while reasoning with the capacity of a 30B one — built for hard, practical work : writing and debugging real code, driving tools and agents, designing production grade interfaces, and visual understanding over images and video, inside a single model with a context window of up to 1M tokens . It is the fast, multimodal tier of the Salience family — engineered for people who care less about chat pleasantries and more about whether the model can do the thing : ship the function, find the bug, call the right tool, design the screen, read the diagram. Highlights Bigger and faster at once. Sparse activation means ~3B of compute for 30B of knowledge — roughly 2× the decode speed of a dense 8B at far greater capacity. Code & agentic first. Tuned to produce runnable cod…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy