LFM2.5 8B A1B APEX GGUF APEX (Adaptive Precision for EXpert Models) quantizations of LiquidAI/LFM2.5 8B A1B. Brought to you by the LocalAI team APEX Project Available Files File Profile Size Best For LFM2.5 8B A1B APEX I Quality.gguf I Quality 6.1 GB Highest quality with imatrix LFM2.5 8B A1B APEX Quality.gguf Quality 6.1 GB Highest quality standard LFM2.5 8B A1B APEX I Balanced.gguf I Balanced 6.3 GB Best overall quality/size ratio LFM2.5 8B A1B APEX Balanced.gguf Balanced 6.3 GB General purpose LFM2.5 8B A1B APEX I Compact.gguf I Compact 4.2 GB Consumer GPUs, best quality/size LFM2.5 8B A1B APEX Compact.gguf Compact 4.2 GB Consumer GPUs LFM2.5 8B A1B APEX I Mini.gguf I Mini 3.6 GB Smallest viable, fastest inference (I variants use imatrix calibrated quantization; the matching base profiles are the same size without imatrix weighting.) What is APEX? APEX is a quantization strategy for Mixture of Experts (MoE) models. It classifies tensors by role (routed expert, shared expert, attention, token mixing) and applies a layer wise precision gradient — edge layers get higher precision, middle layers get more aggressive compression. I variants use diverse imatrix calibration (chat, code,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy