Prism ML Website Whitepaper Demo & Examples Discord 1 bit Bonsai 27B Full 27B class reasoning in binary transformer weights — the first 27B class model to run on a phone ~14.2x smaller than FP16 ~90% of FP16 intelligence retained ~11 tok/s on iPhone 17 Pro Max Highlights ~3.9 GB deployed footprint (down from ~54 GB FP16) — fits within the per app memory budget of a high end phone such as the iPhone 17 Pro Max Retains thinking, reasoning, and agentic behavior deep in the sub 4 bit regime, where conventional low bit representations collapse — 76.11 average across 15 thinking mode benchmarks (89.5% of FP16), including math at 91.66 and coding at 81.88 End to end binary language weights across embeddings, attention projections, MLP projections, and LM head, at a true 1.125 bits per weight — no high precision escape hatches behind a low bit label; the vision tower ships in compact 4 bit HQQ 262K token context on device, kept practical by the Qwen3.6 27B hybrid attention backbone (~75% linear attention) and 4 bit KV cache quantization First interactive 27B class generation on a phone : ~11 tok/s on iPhone 17 Pro Max; ~44 tok/s on an Apple M5…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy