Holo 3.1 35B A3B — GGUF Quantizations GGUF quantizations of Hcompany/Holo 3.1 35B A3B for use with llama.cpp. About the Model Architecture: Qwen3.5 MoE (35B total, ~3B active parameters) Type: Vision Language Model (VLM) for computer use agents Context: 262,144 tokens License: Apache 2.0 Note: BF16 GGUF and imatrix sourced from Hcompany/Holo 3.1 35B A3B GGUF. Available Files Filename Quant Bits Size (est.) Holo 3.1 35B A3B Q2 K.gguf Q2 K 2 ~12 GB Holo 3.1 35B A3B Q3 K S.gguf Q3 K S 3 ~15 GB Holo 3.1 35B A3B Q3 K M.gguf Q3 K M 3 ~17 GB Holo 3.1 35B A3B Q3 K L.gguf Q3 K L 3 ~18 GB Holo 3.1 35B A3B IQ3 XXS.gguf IQ3 XXS 3 ~14 GB Holo 3.1 35B A3B IQ3 S.gguf IQ3 S 3 ~15 GB Holo 3.1 35B A3B Q4 K S.gguf Q4 K S 4 ~20 GB Holo 3.1 35B A3B Q4 K M.gguf Q4 K M 4 ~22 GB Holo 3.1 35B A3B Q5 K S.gguf Q5 K S 5 ~24 GB Holo 3.1 35B A3B Q5 K M.gguf Q5 K M 5 ~26 GB Holo 3.1 35B A3B Q6 K.gguf Q6 K 6 ~30 GB Holo 3.1 35B A3B Q8 0.gguf Q8 0 8 ~39 GB mmproj Holo 3.1 35B A3B F16.gguf F16 16 ~0.9 GB BF16 GGUF (69.4 GB) available at source repo. Usage Quantization Details BF16 GGUF source: Hcompany/Holo 3.1 35B A3B GGUF imatrix: Pre computed, sourced from BF16 repo (192 MB) Tool: llama.cpp llama quantize Quanti…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy