MiniCPM5 1B Claude Opus Fable5 V2 Thinking GGUF GGUF quantizations of MiniCPM5 1B Claude Opus Fable5 V2 Thinking for llama.cpp, Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes. 中文说明 This repository provides local deployment builds of a 1B Thinking model fine tuned on Fable 5 data (V2) atop openbmb/MiniCPM5 1B. Compared with V1, V2 strengthens tool calling / function calling , while keeping MiniCPM5's native chat template embedded in the GGUF files. Transformers checkpoint: MiniCPM5 1B Claude Opus Fable5 V2 Thinking Previous GGUF version: MiniCPM5 1B Claude Opus Fable5 Thinking GGUF (V1) Files File Quant Size Notes MiniCPM5 1B Claude Opus Fable5 V2 Thinking Q8 0.gguf Q8 0 ~1.1 GB recommended default MiniCPM5 1B Claude Opus Fable5 V2 Thinking F16.gguf F16 ~2.1 GB full precision conversion base Q8 0 is the recommended default quant for this 1B model. Quick start llama.cpp ( llama cli ) The model supports up to 128K tokens (131,072) per config.json . Set c according to your available VRAM/RAM. llama.cpp server LM Studio / jan / KoboldCpp Load any .gguf file from this repository. The MiniCPM5 chat template is embedded in the GGUF metadata. Sampling recommendations Generation…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy