Llamacpp imatrix Quantizations of UI TARS 72B DPO Using llama.cpp release b4514 for quantization. Original model: https://huggingface.co/bytedance research/UI TARS 72B DPO All quants made using imatrix option with dataset from here Run them in LM Studio Prompt format Download a file (not the whole branch) from below: Filename Quant type File Size Split Description UI TARS 72B DPO Q8 0.gguf Q8 0 77.26GB true Extremely high quality, generally unneeded but max available quant. UI TARS 72B DPO Q6 K.gguf Q6 K 64.35GB true Very high quality, near perfect, recommended . UI TARS 72B DPO Q5 K M.gguf Q5 K M 54.45GB true High quality, recommended . UI TARS 72B DPO Q5 K S.gguf Q5 K S 51.38GB true High quality, recommended . UI TARS 72B DPO Q4 K M.gguf Q4 K M 47.42GB false Good quality, default size for most use cases, recommended . UI TARS 72B DPO Q4 1.gguf Q4 1 45.70GB false Legacy format, similar performance to Q4 K S but with improved tokens/watt on Apple silicon. UI TARS 72B DPO Q4 K S.gguf Q4 K S 43.89GB false Slightly lower quality with more space savings, recommended . UI TARS 72B DPO Q4 0.gguf Q4 0 41.38GB false Legacy format, offers online repacking for ARM and AVX CPU inference. UI T…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy