Llamacpp imatrix Quantizations of InternVL3 5 4B by OpenGVLab Using llama.cpp release b6258 for quantization. Original model: https://huggingface.co/OpenGVLab/InternVL3 5 4B All quants made using imatrix option with dataset from here combined with a subset of combined all small.parquet from Ed Addario here Run them in LM Studio Run them directly with llama.cpp, or any other llama.cpp based project Prompt format No prompt format found, check original model page Download a file (not the whole branch) from below: Filename Quant type File Size Split Description InternVL3 5 4B bf16.gguf bf16 8.83GB false Full BF16 weights. InternVL3 5 4B Q8 0.gguf Q8 0 4.69GB false Extremely high quality, generally unneeded but max available quant. InternVL3 5 4B Q6 K L.gguf Q6 K L 3.81GB false Uses Q8 0 for embed and output weights. Very high quality, near perfect, recommended . InternVL3 5 4B Q6 K.gguf Q6 K 3.63GB false Very high quality, near perfect, recommended . InternVL3 5 4B Q5 K L.gguf Q5 K L 3.40GB false Uses Q8 0 for embed and output weights. High quality, recommended . InternVL3 5 4B Q5 K M.gguf Q5 K M 3.16GB false High quality, recommended . InternVL3 5 4B Q5 K S.gguf Q5 K S 3.09GB false Hi…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy