Prism ML Website Whitepaper Demo & Examples Discord Ternary Bonsai 27B Full 27B class reasoning in ternary transformer weights — on everyday laptops \~9.4x smaller than FP16 (ideal) 95% of FP16 intelligence retained \~26 tok/s on an Apple M5 Pro laptop Highlights \~7.2 GB deployed footprint (down from \~54 GB FP16) — full 27B class reasoning on a standard laptop or a single GPU 95% of FP16 intelligence retained : 80.49 average across 15 thinking mode benchmarks — a higher score than the conventional IQ2 XXS build (72.73) at less than two thirds of its footprint Retains thinking, reasoning, and agentic behavior deep in the sub 4 bit regime, where conventional low bit representations collapse: math within two points of full precision (93.40), coding at 85.96, agentic tool use at 74.01 End to end ternary language weights across embeddings, attention projections, MLP projections, and LM head, at a true 1.71 bits per weight — no high precision escape hatches behind a low bit label; the vision tower ships in compact 4 bit HQQ 262K token context on device, kept practical by the Qwen3.6 27B hybrid attention backbone (\~75% linear attention) and…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy