Note: This model is no longer the optimal W8A8 quantization, please consider using a better quantization model I made later: noneUsername/Mistral Nemo Instruct 2407 W8A8 Dynamic Per Token better My first quantization uses the quantization method provided by vllm: https://docs.vllm.ai/en/latest/quantization/int8.html NUM CALIBRATION SAMPLES = 2048 MAX SEQUENCE LENGTH = 8192 smoothing strength=0.8 I will verify the validity of the model and update the readme as soon as possible. edit: The performance in my ERP test was comparable to Mistral Nemo Instruct 2407 GPTQ INT8, which I consider a successful quantization. vllm (pretrained=/root/autodl tmp/Mistral Nemo Instruct 2407,add bos token=true,tensor parallel size=2,max model len=4096,gpu memory utilization=0.85,swap space=0), gen kwargs: (None), limit: 250.0, num fewshot: 5, batch size: auto Tasks Version Filter n shot Metric Value Stderr : : : : gsm8k 3 flexible extract 5 exact match ↑ 0.800 ± 0.0253 strict match 5 exact match ↑ 0.784 ± 0.0261 lm eval model vllm model args pretrained="/mnt/e/Code/models/Mistral Nemo Instruct 2407 W8A8 Dynamic Per Token",add bos token=true,dtype=half,tensor parallel size=2,max model len=4096,gpu memor…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy