Llamacpp imatrix Quantizations of GLM 4.5 Air REAP 82B A12B by cerebras Using llama.cpp release b6818 for quantization. Original model: https://huggingface.co/cerebras/GLM 4.5 Air REAP 82B A12B All quants made using imatrix option with dataset from here combined with a subset of combined all small.parquet from Ed Addario here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format No chat template specified so default is used. This may be incorrect, check original model card for details. Download a file (not the whole branch) from below: Filename Quant type File Size Split Description GLM 4.5 Air REAP 82B A12B Q8 0.gguf Q8 0 90.37GB true Extremely high quality, generally unneeded but max available quant. GLM 4.5 Air REAP 82B A12B Q6 K.gguf Q6 K 76.21GB true Very high quality, near perfect, recommended . GLM 4.5 Air REAP 82B A12B Q5 K M.gguf Q5 K M 64.39GB true High quality, recommended . GLM 4.5 Air REAP 82B A12B Q5 K S.gguf Q5 K S 60.49GB true High quality, recommended . GLM 4.5 Air REAP 82B A12B Q4 K L.gguf Q4 K L 57.03GB true Uses Q8 0 for embed and output weights. Good quality, recommended . GLM 4.5 Air REAP 82B A12B Q4 K M.ggu…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy