Prism ML Website Whitepaper Demo & Examples Discord 1 bit Bonsai 27B — GGUF Full 27B class reasoning in binary transformer weights, for llama.cpp (CUDA, Metal, CPU) \~14.2x smaller than FP16 \~90% of FP16 intelligence retained \~44 tok/s on an Apple M5 Pro laptop Highlights \~3.9 GB deployed footprint (down from \~54 GB FP16) — a 27B model on everyday laptops and single GPUs Retains thinking, reasoning, and agentic behavior deep in the sub 4 bit regime, where conventional low bit representations collapse — 76.11 average across 15 thinking mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 End to end binary language weights across embeddings, attention projections, MLP projections, and LM head, at a true 1.125 bits per weight — no high precision escape hatches behind a low bit label; the vision tower ships in compact 4 bit HQQ 262K token context on device, kept practical by the Qwen3.6 27B hybrid attention backbone (\~75% linear attention) and 4 bit KV cache quantization GGUF Q1 0 g128 format with custom 1 bit hybrid attention kernels for llama.cpp (CUDA, Metal) — packed weights are consumed directly, never expa…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy