openPangu 2.0 Flash GGUF GGUF conversion of openPangu 2.0 Flash (92B MoE, \~6B active parameters, 512K context), converted directly from the original bf16 safetensors. [!IMPORTANT] These files require a llama.cpp fork with openPangu support : https://github.com/mrexodia/llama.cpp openPangu 2.0 Flash Upstream llama.cpp cannot load this architecture yet. Supported by the fork: MLA attention, DSA sparse attention (lightning indexer top 2048) on the global layers, per layer sliding window attention, manifold hyper connections (mHC), MoME convolutions, learned attention sinks, tool calling + reasoning parsing, and optional multi token prediction (MTP) self speculative decoding. Files File Size Notes openPangu 2.0 Flash base Q3 K M.gguf 42 GB fits 64 GB Apple Silicon openPangu 2.0 Flash base Q4 K M.gguf 52 GB recommended for speed openPangu 2.0 Flash base Q8 0.gguf 91 GB recommended for quality (fits DGX Spark) openPangu 2.0 Flash base BF16.gguf 183 GB requant source openPangu 2.0 Flash mtp Q8 0.gguf 9.2 GB optional MTP draft head openPangu 2.0 Flash mtp BF16.gguf 19 GB requant source The base files omit the 3 MTP (NextN) layers; the mtp files contain only them, for use as a speculative…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy