Prism ML Website Whitepaper Demo & Examples Colab Notebook Discord Bonsai 8B GGUF 1bit End to end 1 bit language model for llama.cpp (CUDA, Metal, CPU) 14.1x smaller than FP16 6.2x faster on RTX 4090 4 5x lower energy/token Highlights 1.15 GB parameter memory (down from 16.38 GB FP16) — fits on virtually any device with a GPU End to end 1 bit weights across embeddings, attention projections, MLP projections, and LM head GGUF Q1 0 (g128) format with inline dequantization kernels — no FP16 materialization Cross platform : CUDA (RTX/datacenter), Metal (Mac), Android, CPU Competitive benchmarks : 70.5 avg score across 6 categories, matching full precision 8B models at 1/14th the size MLX companion : also available as MLX 1 bit g128 for native Apple Silicon inference Resources Google Colab — try Bonsai in your browser, no setup required Whitepaper — for more details on Bonsai, check out our whitepaper Demo repo — comprehensive examples for serving, benchmarking, and integrating Bonsai Discord — join the community for support, discussion, and updates 1 bit kernels : llama.cpp fork (CUDA + Metal) · MLX fork (Apple Silicon) · mlx…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy