Inkling MXFP4 Model Overview Model Architecture: Thinking Machines Lab Inkling Input: Text, Image, Audio Output: Text Inference Engine: TokenSpeed Model Optimizer: AMD Quark (0.12.post1+rocm72.torch2.11) Quantized layers: MoE routed experts only Weight quantization: OCP MXFP4, static Activation quantization: OCP MXFP4, dynamic This model was built by applying AMD Quark MXFP4 quantization to the BF16 Thinking Machines Lab Inkling checkpoint. The quantization targets the MoE routed experts, while attention layers and shared experts are kept in BF16. Environment The quantization workflow was prepared on an AMD gfx950 system. The inspected container environment was: GPU: AMD MI350/MI355 Target graphics version: gfx950 ROCm: 7.2.1 amdgpu driver: 6.16.13 OS: Linux 6.8.0 84, x86 64 Python: 3.12.3 PyTorch: 2.13.0+rocm7.1 AMD Quark: 0.12.post1+rocm72.torch2.11 Safetensors: 0.8.0 Transformers: 5.13.1 Create and activate the Quark environment: Install the required packages: Model Quantization The model was quantized with the Quark file to file flow. This avoids loading the full BF16 checkpoint into GPU memory at once, which is important for very large MoE checkpoints. Run the quantization scr…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy