Importance Matrix Calibration Datasets
This repository provides calibration datasets used to generate importance matrices (imatrix), which are required to minimize errors when quantizing models with LLaMA C++.
The llama-imatrix program cannot handle parquet files directly and thus requires them to be converted into text format first. There are many ways to do this but a simple approach is to use DuckDB with the following command: duckdb -noheader -ascii -c "SELECT content FROM 'tools_micro.parquet';" > tools_micro.txt
Code calibration datasets
This dataset consists of cleaned and de-duplicated code prompts and is available in six sizes, ranging from huge (~ 200,000 lines equivalent to approx. 14.5M tokens), to micro (~ 6,200 lines and 2.4M tokens avg).
Original data sourced from Vezora/Open-Critic-GPT, OpenCoder-LLM/opc-sft-stage2, ise-uiuc/Magicoder-Evol-Instruct-110K, and Multilingual-Multimodal-NLP/McEval-Instruct
Math calibration datasets
This dataset consists of cleaned and de-duplicated math prompts and is available in six sizes, ranging from huge (~ 200,000 lines equivalent to approx. 6 million), to micro (~ 6,250 lines and 0.9 million tokens avg).
Original data sourced from nvidia/OpenMathInstruct-2
Tools calibration datasets
This dataset consists of cleaned and de-duplicated tool prompts and is available in six sizes, ranging from huge (~ 100,000 lines equivalent to approx. 10 million tokens), to micro (~ 3,100 lines and 1 million tokens).
Original data sourced from BitAgent/tool_calling and JungHun/Efficient_ToolCalling
Language calibration datasets
This dataset consists of cleaned and de-duplicated text prompts for 18 different languages. Each language file is available in five sizes, ranging from large (~ 25,000 lines equivalent to approx. 725K tokens), to micro (~ 1,600 lines and 125K tokens avg).
Original data sourced from HuggingFaceFW/fineweb, HuggingFaceFW/fineweb-2, and Common Crawl
| File | Language | Lines |
|---|
| text_ar_large | Arabic | 25,000 |
| text_ar_medium | Arabic | 12,500 |
| text_ar_small | Arabic | 6,250 |
| text_ar_tiny | Arabic | 3,125 |
| text_ar_micro | Arabic | 1,562 |
| text_cn_large | Chinese | 25,000 |
| text_cn_medium | Chinese | 12,500 |
| text_cn_small | Chinese | 6,250 |
| text_cn_tiny | Chinese | 3,125 |
| text_cn_micro | Chinese | 1,562 |
| text_de_large | German | 25,000 |
| text_de_medium | German | 12,500 |
| text_de_small | German | 6,250 |
| text_de_tiny | German | 3,125 |
| text_de_micro | German | 1,562 |
| text_en_large | English | 25,000 |
| text_en_medium | English | 12,500 |
| text_en_small | English | 6,250 |
| text_en_tiny | English | 3,125 |
| text_en_micro | English | 1,562 |
| text_es_large | Spanish | 25,000 |
| text_es_medium | Spanish | 12,500 |
| text_es_small | Spanish | 6,250 |
| text_es_tiny | Spanish | 3,125 |
| text_es_micro | Spanish | 1,562 |
| text_fr_large | French | 25,000 |
| text_fr_medium | French | 12,500 |
| text_fr_small | French | 6,250 |
| text_fr_tiny | French | 3,125 |
| text_fr_micro | French | 1,562 |
| text_hi_large | Hindi | 25,000 |
| text_hi_medium | Hindi | 12,500 |
| text_hi_small | Hindi | 6,250 |
| text_hi_tiny | Hindi | 3,125 |
| text_hi_micro | Hindi | 1,562 |
| text_id_large | Indonesian | 24,999 |
| text_id_medium | Indonesian | 12,500 |
| text_id_small | Indonesian | 6,250 |
| text_id_tiny | Indonesian | 3,125 |
| text_id_micro | Indonesian | 1,562 |
| text_it_large | Italian | 25,000 |
| text_it_medium | Italian | 12,500 |
| text_it_small | Italian | 6,250 |
| text_it_tiny | Italian | 3,125 |
| text_it_micro | Italian | 1,562 |
| text_jp_large | Japanese | 25,000 |
| text_jp_medium | Japanese | 12,500 |
| text_jp_small | Japanese | 6,250 |
| text_jp_tiny | Japanese | 3,125 |
| text_jp_micro | Japanese | 1,562 |
| text_mm_large | Burmese | 25,000 |
| text_mm_medium | Burmese | 12,500 |
| text_mm_small | Burmese | 6,250 |
| text_mm_tiny | Burmese | 3,125 |
| text_mm_micro | Burmese | 1,562 |
| text_nl_large | Dutch | 25,000 |
| text_nl_medium | Dutch | 12,500 |
| text_nl_small | Dutch | 6,250 |
| text_nl_tiny | Dutch | 3,125 |
| text_nl_micro | Dutch | 1,562 |
| text_ph_large | Filipino | 25,000 |
| text_ph_medium | Filipino | 12,500 |
| text_ph_small | Filipino | 6,250 |
| text_ph_tiny | Filipino | 3,125 |
| text_ph_micro | Filipino | 1,562 |
| text_pl_large | Polish | 25,000 |
| text_pl_medium | Polish | 12,500 |
| text_pl_small | Polish | 6,250 |
| text_pl_tiny | Polish | 3,125 |
| text_pl_micro | Polish | 1,562 |
| text_pt_large | Portuguese | 25,000 |
| text_pt_medium | Portuguese | 12,500 |
| text_pt_small | Portuguese | 6,250 |
| text_pt_tiny | Portuguese | 3,125 |
| text_pt_micro | Portuguese | 1,562 |
| text_ru_large | Russian | 25,000 |
| text_ru_medium | Russian | 12,500 |
| text_ru_small | Russian | 6,250 |
| text_ru_tiny | Russian | 3,125 |
| text_ru_micro | Russian | 1,562 |
| text_th_large | Thai | 25,000 |
| text_th_medium | Thai | 12,500 |
| text_th_small | Thai | 6,250 |
| text_th_tiny | Thai | 3,125 |
| text_th_micro | Thai | 1,562 |
| text_vn_large | Vietnamese | 25,000 |
| text_vn_medium | Vietnamese | 12,500 |
| text_vn_small | Vietnamese | 6,250 |
| text_vn_tiny | Vietnamese | 3,125 |
| text_vn_micro | Vietnamese | 1,562 |
Language groups
In addition to single language files, the dataset includes randomized and files by language family/region and all languages in dataset
All languages (all)
European languages: English, French, German, Italian, Portuguese & Spanish (eur)
Germanic languages: Dutch, English & German (gem)
Romance languages: French, Italian, Portuguese & Spanish (roa)
Rest of World: Arabic, Chinese, Hindi & Japanese (row)
Southeast Asia languages: Burmese, Filipino, Indonesian, Thai & Vietnamese (sea)
Slavic languages: Polish & Russian (sla)
Math & Code calibration datasets
This dataset combines math and code prompts into single calibration files.
Tool, Math, Code and Language calibration datasets
This dataset combines tool, math, code and language prompts into single calibration files.
| File | Language | Lines |
|---|
| combined_ar_huge | Arabic | 100,000 |
| combined_ar_large | Arabic | 50,000 |
| combined_ar_medium | Arabic | 25,000 |
| combined_ar_small | Arabic | 12,500 |
| combined_ar_tiny | Arabic | 6,248 |
| combined_ar_micro | Arabic | 3,124 |
| combined_cn_huge | Chinese | 100,000 |
| combined_cn_large | Chinese | 50,000 |
| combined_cn_medium | Chinese | 25,000 |
| combined_cn_small | Chinese | 12,500 |
| combined_cn_tiny | Chinese | 6,248 |
| combined_cn_micro | Chinese | 3,124 |
| combined_de_huge | German | 100,000 |
| combined_de_large | German | 50,000 |
| combined_de_medium | German | 25,000 |
| combined_de_small | German | 12,500 |
| combined_de_tiny | German | 6,248 |
| combined_de_micro | German | 3,124 |
| combined_en_huge | English | 99,999 |
| combined_en_large | English | 50,000 |
| combined_en_medium | English | 25,000 |
| combined_en_small | English | 12,500 |
| combined_en_tiny | English | 6,248 |
| combined_en_micro | English | 3,124 |
| combined_es_huge | Spanish | 100,000 |
| combined_es_large | Spanish | 50,000 |
| combined_es_medium | Spanish | 25,000 |
| combined_es_small | Spanish | 12,500 |
| combined_es_tiny | Spanish | 6,248 |
| combined_es_micro | Spanish | 3,124 |
| combined_fr_huge | French | 100,000 |
| combined_fr_large | French | 50,000 |
| combined_fr_medium | French | 25,000 |
| combined_fr_small | French | 12,500 |
| combined_fr_tiny | French | 6,248 |
| combined_fr_micro | French | 3,124 |
| combined_hi_huge | Hindi | 100,000 |
| combined_hi_large | Hindi | 50,000 |
| combined_hi_medium | Hindi | 25,000 |
| combined_hi_small | Hindi | 12,500 |
| combined_hi_tiny | Hindi | 6,248 |
| combined_hi_micro | Hindi | 3,124 |
| combined_id_huge | Indonesian | 99,999 |
| combined_id_large | Indonesian | 50,000 |
| combined_id_medium | Indonesian | 25,000 |
| combined_id_small | Indonesian | 12,500 |
| combined_id_tiny | Indonesian | 6,248 |
| combined_id_micro | Indonesian | 3,124 |
| combined_it_huge | Italian | 100,000 |
| combined_it_large | Italian | 50,000 |
| combined_it_medium | Italian | 25,000 |
| combined_it_small | Italian | 12,500 |
| combined_it_tiny | Italian | 6,248 |
| combined_it_micro | Italian | 3,124 |
| combined_jp_huge | Japanese | 100,000 |
| combined_jp_large | Japanese | 50,000 |
| combined_jp_medium | Japanese | 25,000 |
| combined_jp_small | Japanese | 12,500 |
| combined_jp_tiny | Japanese | 6,248 |
| combined_jp_micro | Japanese | 3,124 |
| combined_mm_huge | Burmese | 100,000 |
| combined_mm_large | Burmese | 50,000 |
| combined_mm_medium | Burmese | 25,000 |
| combined_mm_small | Burmese | 12,500 |
| combined_mm_tiny | Burmese | 6,248 |
| combined_mm_micro | Burmese | 3,124 |
| combined_nl_huge | Dutch | 100,000 |
| combined_nl_large | Dutch | 50,000 |
| combined_nl_medium | Dutch | 25,000 |
| combined_nl_small | Dutch | 12,500 |
| combined_nl_tiny | Dutch | 6,248 |
| combined_nl_micro | Dutch | 3,124 |
| combined_ph_huge | Filipino | 100,000 |
| combined_ph_large | Filipino | 49,999 |
| combined_ph_medium | Filipino | 25,000 |
| combined_ph_small | Filipino | 12,500 |
| combined_ph_tiny | Filipino | 6,248 |
| combined_ph_micro | Filipino | 3,124 |
| combined_pl_huge | Polish | 100,000 |
| combined_pl_large | Polish | 50,000 |
| combined_pl_medium | Polish | 25,000 |
| combined_pl_small | Polish | 12,500 |
| combined_pl_tiny | Polish | 6,248 |
| combined_pl_micro | Polish | 3,124 |
| combined_pt_huge | Portuguese | 100,000 |
| combined_pt_large | Portuguese | 50,000 |
| combined_pt_medium | Portuguese | 25,000 |
| combined_pt_small | Portuguese | 12,500 |
| combined_pt_tiny | Portuguese | 6,248 |
| combined_pt_micro | Portuguese | 3,124 |
| combined_ru_huge | Russian | 99,999 |
| combined_ru_large | Russian | 50,000 |
| combined_ru_medium | Russian | 25,000 |
| combined_ru_small | Russian | 12,500 |
| combined_ru_tiny | Russian | 6,248 |
| combined_ru_micro | Russian | 3,124 |
| combined_th_huge | Thai | 100,000 |
| combined_th_large | Thai | 50,000 |
| combined_th_medium | Thai | 25,000 |
| combined_th_small | Thai | 12,500 |
| combined_th_tiny | Thai | 6,248 |
| combined_th_micro | Thai | 3,124 |
| combined_vn_huge | Vietnamese | 99,999 |
| combined_vn_large | Vietnamese | 50,000 |
| combined_vn_medium | Vietnamese | 25,000 |
| combined_vn_small | Vietnamese | 12,499 |
| combined_vn_tiny | Vietnamese | 6,248 |
| combined_vn_micro | Vietnamese | 3,124 |
Tool, Math, Code and Language groups calibration datasets
In addition to single tool, math, code and language files, the dataset includes combined and randomized files by language family/region and all languages in dataset
All languages (all)
European languages: English, French, German, Italian, Portuguese & Spanish (eur)
Germanic languages: Dutch, English & German (gem)
Romance languages: French, Italian, Portuguese & Spanish (roa)
Rest of World: Arabic, Chinese, Hindi & Japanese (row)
Southeast Asia languages: Burmese, Filipino, Indonesian, Thai & Vietnamese (sea)
Slavic languages: Polish & Russian (sla)