GLM 4 32B 0414 Quantized with GPTQ (4 Bit weight only, W4A16) This repo contains GLM 4 32B 0414 quantized with asymmetric GPTQ to 4 bit to make it suitable for consumer hardware. The model was calibrated with 2048 samples of max sequence length 4096 from the dataset mit han lab/pile val backup . This is my very first quantized model, I welcome suggestions. 2048/4096 were chosen over the default of 512/2048 to minimize overfitting risk and maximize convergence. They also happen to fit in my GPU. Original Model: zai org/GLM 4 32B 0414 📥 Usage & Running Instructions The model was tested with vLLM, here is a script suitable for 32GB VRAM GPUs. 🔬 Quantization method The llmcompressor library was used with the following recipe for asymmetric GPTQ: and calibrated on 2048 samples, 4096 sequence length of mit han lab/pile val backup
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy