Nanbeige4.2 3B GGUF Quants This repository contains GGUF quantizations for Nanbeige/Nanbeige4.2 3B . Original Model: Nanbeige/Nanbeige4.2 3B Base Architecture: Looped Transformer (3B non embedding parameters) Quantization Format: GGUF ( Q4 K M , Q4 K S , Q5 K M , Q6 K , Q8 0 ) Available Files & Quantization Details File Name Size Quant Method Description Nanbeige4.2 3B Q4 K M.gguf ~2.57 GB Q4 K M 4 bit medium. Recommended balance of speed, memory usage, and quality. Nanbeige4.2 3B Q4 K S.gguf ~2.50 GB Q4 K S 4 bit small. Slightly lower memory footprint. Nanbeige4.2 3B Q5 K M.gguf ~2.99 GB Q5 K M 5 bit medium. Higher precision with slight increase in size. Nanbeige4.2 3B Q6 K.gguf ~3.42 GB Q6 K 6 bit quantization. Very close to FP16 performance. Nanbeige4.2 3B Q8 0.gguf ~4.43 GB Q8 0 8 bit quantization. Maximum quality for GGUF. Usage Guide 1. Running with llama.cpp For full support, clone the official or nanbeige42 fork of llama.cpp : bash Clone the repository with Nanbeige support git clone b nanbeige42 https://github.com/Nanbeige/llama.cpp.git cd llama.cpp Build with CUDA support cmake B build DGGML CUDA=ON cmake build build config Release j Download a model from this repository…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy