Granite 4.0 h tiny FP8 dynamic Model Overview Model Architecture: GraniteMoeHybridForCausalLM Input: Text Output: Text Model Optimizations: Weight quantization: FP8 Activation quantization: FP8 Release Date: Version: 1.0 Model Developers: : Red Hat ModelCar Storage URI : oci://registry.redhat.io/rhai/modelcar granite 4 0 h tiny fp8 dynamic:3.0 Validated on vLLM: 0.13.0 Validated on RHAIIS: 3.3 Validated on RHOAI: 3.3 Quantized version of ibm granite/granite 4.0 h tiny. Model Optimizations This model was obtained by quantizing the weights and activations of ibm granite/granite 4.0 h tiny to FP8 data type. This optimization reduces the number of bits per parameter from 16 to 8, reducing the disk size and GPU memory requirements by approximately 50%. Only the weights and activations of the linear operators within transformers blocks of the language model are quantized. Deployment Use with vLLM 1. Install vLLM from main: 2. Initialize vLLM server: 3. Send requests to the server: Creation This model was quantized using the llm compressor library as shown below. Creation details Install specific llm compression version: Evaluation The model was evaluated on the OpenLLM leaderboard task,…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy