Try LFM • Docs • LEAP • Discord LFM2 24B A2B LFM2 is a family of hybrid models designed for on device deployment. LFM2 24B A2B is the largest model in the family, scaling the architecture to 24 billion parameters while keeping inference efficient. Best in class efficiency : A 24B MoE model with only 2B active parameters per token, fitting in 32 GB of RAM for deployment on consumer laptops and desktops. Fast edge inference : 112 tok/s decode on AMD CPU, 293 tok/s on H100. Fits in 32B GB of RAM with day one support llama.cpp, vLLM, and SGLang. Predictable scaling : Quality improves log linearly from 350M to 24B total parameters, confirming the LFM2 hybrid architecture scales reliably across nearly two orders of magnitude. Find more information about LFM2 24B A2B in our blog post. 🗒️ Model Details LFM2 24B A2B is a general purpose instruct model (without reasoning traces) with the following features: Property LFM2 8B A1B LFM2 24B A2B Total parameters 8.3B 24B Active parameters 1.5B 2.3B Layers 24 (18 conv + 6 attn) 40 (30 conv + 10 attn) Context length 32,768 tokens 32,768 tokens Vocabulary size 65,536 65,536 Training precision Mixed BF16/FP8 Mixed BF16/FP8 Training budget 12 trillio…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy