ZAYA1 8B GGUF Quantizations This repository contains GGUF quantizations of Zyphra/ZAYA1 8B. ⚠️ CRITICAL USAGE NOTE: The zaya architecture (which utilizes a unique Compressed Convolutional Attention mechanism and an MLP router) is currently undergoing experimental integration into llama.cpp . These quantizations were generated using the bleeding edge Draft PR 23112. To protect the model's complex reasoning and routing logic, the highly sensitive cca conv grp layers were explicitly excluded from quantization and remain in higher precision. To run these models, you must compile llama.cpp locally from that specific PR branch until official support is merged into the master branch. Model Details ZAYA1 8B is a frontier level reasoning Mixture of Experts (MoE) model designed for high intelligence density and local deployment. Total Parameters: ~8.4B Active Parameters (per token): ~760M Architecture: zaya (Sparse MoE) License: Apache 2.0 Creator: Zyphra Available Quants Format File Size Description : : : Q3 K M 4.51 GB Smallest viable quant. Heavy compression, potential logic degradation. Q4 K S 5.26 GB Very small. High compression, suitable for strict VRAM limits. Q4 K M 5.57 GB Recommend…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy