Mistral Small 3.2 24B Instruct 2506 NVFP4 Model Overview Model Architecture: unsloth/Mistral Small 3.2 24B Instruct 2506 Input: Text Output: Text Model Optimizations: Weight quantization: FP4 Activation quantization: FP4 Out of scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages other than English. Release Date: 10/29/2025 Version: 1.0 Model Developers: RedHatAI This model is a quantized version of unsloth/Mistral Small 3.2 24B Instruct 2506. It was evaluated on a several tasks to assess the its quality in comparison to the unquatized model. Model Optimizations This model was obtained by quantizing the weights and activations of unsloth/Mistral Small 3.2 24B Instruct 2506 to FP4 data type, ready for inference with vLLM =0.9.1 This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%. Only the weights and activations of the linear operators within transformers blocks are quantized using LLM Compressor. Deployment Use with vLLM 1. Initialize vLLM server: 2. Send requests to the server: Creation This model was created by applying…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy