Prism ML Website Whitepaper Demo & Examples Colab Notebook Discord Bonsai 8B mlx 1bit End to end 1 bit language model for Apple Silicon 12.8x smaller than FP16 8.4x faster on M4 Pro 44 tok/s on iPhone runs on Mac, iPhone, iPad Highlights 1.28 GB parameter memory (down from 16.38 GB FP16) — runs comfortably on any Mac or iPhone End to end 1 bit weights across embeddings, attention projections, MLP projections, and LM head MLX native format (1 bit g128) with inline dequantization kernels — no FP16 materialization Competitive benchmarks : 70.5 avg score across 6 categories, matching full precision 8B models at 1/14th the size Cross platform companion : also available as GGUF Q1 0 g128 for llama.cpp Resources Google Colab — try Bonsai in your browser, no setup required Whitepaper — for more details on Bonsai, check out our whitepaper Demo repo — comprehensive examples for serving, benchmarking, and integrating Bonsai Discord — join the community for support, discussion, and updates 1 bit kernels : MLX fork (Apple Silicon) · mlx swift fork (iOS/macOS) · llama.cpp fork (CUDA + Metal) Locally AI — we have partnered with Locally A…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy