Mistral Small 3.1 24B Instruct 2503 quantized.w4a16 Model Overview Model Architecture: Mistral3ForConditionalGeneration Input: Text / Image Output: Text Model Optimizations: Weight quantization: INT4 Intended Use Cases: It is ideal for: Fast response conversational agents. Low latency function calling. Subject matter experts via fine tuning. Local inference for hobbyists and organizations handling sensitive data. Programming and math reasoning. Long document understanding. Visual understanding. Out of scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages not officially supported by the model. Release Date: 04/15/2025 Version: 1.0 Validated on: RHOAI 2.20, RHAIIS 3.0, RHELAI 1.5 Model Developers: Red Hat (Neural Magic) Model Optimizations This model was obtained by quantizing the weights of Mistral Small 3.1 24B Instruct 2503 to INT4 data type. This optimization reduces the number of bits per parameter from 16 to 4, reducing the disk size and GPU memory requirements by approximately 75%. Only the weights of the linear operators within transformers blocks are quantized. Weights are quantized using a symmetric per gro…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy