Chat & support: TheBloke's Discord server Want to contribute? TheBloke's Patreon page TheBloke's LLM work is generously supported by a grant from andreessen horowitz (a16z) Dolphin 2.5 Mixtral 8X7B GGUF Model creator: Eric Hartford Original model: Dolphin 2.5 Mixtral 8X7B Description This repo contains GGUF format model files for Eric Hartford's Dolphin 2.5 Mixtral 8X7B. About GGUF GGUF is a new format introduced by the llama.cpp team on August 21st 2023. It is a replacement for GGML, which is no longer supported by llama.cpp. Mixtral GGUF Support for Mixtral was merged into Llama.cpp on December 13th. These Mixtral GGUFs are known to work in: llama.cpp as of December 13th KoboldCpp 1.52 as later LM Studio 0.2.9 and later llama cpp python 0.2.23 and later Other clients/libraries, not listed above, may not yet work. Repositories available GPTQ models for GPU inference, with multiple quantisation parameter options. 2, 3, 4, 5, 6 and 8 bit GGUF models for CPU+GPU inference Eric Hartford's original unquantised fp16 model in pytorch format, for GPU inference and for further conversions Prompt template: ChatML Compatibility These Mixtral GGUFs are compatible with llama.cpp from December…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy