Mistral Small 3.1 24B Instruct 2503 FP8 dynamic Model Overview Model Architecture: Mistral3ForConditionalGeneration Input: Text / Image Output: Text Model Optimizations: Activation quantization: FP8 Weight quantization: FP8 Intended Use Cases: It is ideal for: Fast response conversational agents. Low latency function calling. Subject matter experts via fine tuning. Local inference for hobbyists and organizations handling sensitive data. Programming and math reasoning. Long document understanding. Visual understanding. Out of scope: Use in any manner that violates applicable laws or regulations (including trade compliance laws). Use in languages not officially supported by the model. Release Date: 04/15/2025 Version: 1.0 Validated on: RHOAI 2.20, RHAIIS 3.0, RHELAI 1.5 Model Developers: RedHat (Neural Magic) Model Optimizations This model was obtained by quantizing activations and weights of Mistral Small 3.1 24B Instruct 2503 to FP8 data type. This optimization reduces the number of bits used to represent weights and activations from 16 to 8, reducing GPU memory requirements (by approximately 50%) and increasing matrix multiply compute throughput (by approximately 2x). Weight quant…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy