Granite 4.0 H Small AWQ INT4 Model Details Quantization Details Quantization method: AWQ Bits: 4 Group Size: 32 Calibration Dataset: nvidia/Llama Nemotron Post Training Dataset Quantization Tool: llm compressor Get Started Prerequisite Basic Usage Additional Information Known Issues The model can not load with tensor parallelism and pipeline parallelism. Changelog v1.0.0 Initial quantized model release Granite 4.0 H Small Model Summary: Granite 4.0 H Small is a 32B parameter long context instruct model finetuned from Granite 4.0 H Small Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging. Granite 4.0 instruct models feature improved instruction following (IF) and tool calling capabilities, making them more effective in enterprise applications. Developers: Granite Team, IBM HF Collection: Granite 4.0 Language Models HF Collection GitHub Repository: ibm granite/granite 4.0 language models Website : Granite Docs Release Date…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy