Prism ML Website White Paper Demo & Examples Discord Ternary Bonsai 8B mlx 2bit Ternary (1.58 bit) language model for Apple Silicon 7.1x smaller than FP16 5.2x faster on M4 Pro 27 tok/s on iPhone runs on Mac, iPhone, iPad Highlights 2.15 GiB (2.30 GB) packed 2 bit size (down from 16.38 GB FP16) — runs comfortably on any Mac or iPhone Ternary weights { 1, 0, +1} across embeddings, attention projections, MLP projections, and LM head 75.5 avg benchmark score across 6 categories — competitive with full precision 8B models at 1/9th the size 5 point improvement over our earlier 1 bit Bonsai 8B (70.5) at only ~0.6 GB additional footprint MLX native format with group size 128 and FP16 scaling Resources White Paper Demo repo — examples for serving, benchmarking, and integrating Bonsai Discord — community support and updates Kernels : MLX (Apple Silicon) · mlx swift (iOS/macOS) — 2 bit format is supported out of the box Model Overview Item Specification : : Base model Qwen3 8B Parameters 8.19B (~6.95B non embedding) Architecture GQA (32 query / 8 KV heads), SwiGLU MLP, RoPE, RMSNorm Layers 36 Transformer decoder blocks Context length 65,536 token…
We use cookies for essential functionality and analytics. You can accept or reject analytics cookies.Cookie policy