Ornith 1.0 35B GGUF — iMatrix GGUF GGUF quantizations of deepreinforce ai/Ornith 1.0 35B GGUF, published by Liodon AI. Quick Start llama.cpp Ollama LM Studio / Jan — search liodon ai/Ornith 1.0 35B GGUF imatrix GGUF and pick your quant. Quants Quant Size VRAM est. Notes IQ2 M 11.66 GB ~13 GB 2 bit, iMatrix — smallest usable IQ3 M 15.44 GB ~18 GB 3 bit, iMatrix — great quality/size tradeoff IQ4 XS 18.73 GB ~22 GB 4 bit extra small, iMatrix Q4 K M 21.17 GB ~24 GB 4 bit, iMatrix calibrated (recommended) Q5 K M 24.73 GB ~28 GB 5 bit, iMatrix calibrated Q6 K 28.51 GB ~33 GB 6 bit, iMatrix calibrated, near lossless Q8 0 36.90 GB ~42 GB 8 bit, essentially lossless What is iMatrix? Standard quantization treats all weights equally. iMatrix runs 128 calibration chunks through the full precision model to find which weights matter most, then allocates more precision where it counts. At Q2/Q3/Q4 this means noticeably better coherence and instruction following — same file size, better output . Calibration: 2M tokens of WikiText 103. Also see plain (non iMatrix) quants: liodon ai/Ornith 1.0 35B GGUF GGUF Source Model : deepreinforce ai/Ornith 1.0 35B GGUF License : other Quantized by Liodon AI
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy