Model Overview Model Architecture: Qwen3 30B A3B Thinking 2507 Input: Text Output: Text Supported Hardware Microarchitecture: AMD MI350/MI355 ROCm : 7.0 Operating System(s): Linux Inference Engine: vLLM Model Optimizer: AMD Quark Weight quantization: Perchannel, FP8E4M3, Static Activation quantization: Pertoken, FP8E4M3, Dynamic Calibration Dataset: Pile This model was built with Qwen3 30B A3B Thinking 2507 model by applying AMD Quark for ptpc quantization. Model Quantization The model was quantized from Qwen/Qwen3 30B A3B Thinking 2507 using AMD Quark. The weights are quantized to FP8 and activations are quantized to FP8. Quantization scripts: Accuracy Benchmark Qwen3 30B A3B Thinking 2507 Qwen3 30B A3B Thinking 2507 ptpc(this model) GSM8K 0.755 0.720 Reproduction Docker: rocm/vllm private:rocm7.1 ubuntu22.04 vllm0.11.2 ptpc fp8 The result of GSM8K was obtained using vLLM. vllm version: main(0b2549) aiter version: 0.13.20191203 GSM8K Deployment Use with vLLM This model can be deployed efficiently using the vLLM backend. Evaluation The evaluation results and reproduction script are being prepared. License Modifications Copyright(c) 2025 Advanced Micro Devices, Inc. All rights reser…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy