Try LFM • Docs • LEAP • Discord LFM2‑VL 450M LFM2‑VL is Liquid AI's first series of multimodal models, designed to process text and images with variable resolutions. Built on the LFM2 backbone, it is optimized for low latency and edge AI applications. We're releasing the weights of two post trained checkpoints with 450M (for highly constrained devices) and 1.6B (more capable yet still lightweight) parameters. 2× faster inference speed on GPUs compared to existing VLMs while maintaining competitive accuracy Flexible architecture with user tunable speed quality tradeoffs at inference time Native resolution processing up to 512×512 with intelligent patch based handling for larger images, avoiding upscaling and distortion Find more about our vision language model in the LFM2 VL post and its language backbone in the LFM2 blog post. 📄 Model details Due to their small size, we recommend fine tuning LFM2 VL models on narrow use cases to maximize performance. They were trained for instruction following and lightweight agentic flows. Not intended for safety‑critical decisions. Property LFM2 VL 450M LFM2 VL 1.6B : : Parameters (LM only) 350M 1.2B Vision encoder SigLIP2 NaFlex base (86M) SigL…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy