Release2026-07-14

PrismML released Bonsai 27B on July 14, 2026, a multimodal model compressed from Qwen3.6 27B and shipped in two variants: Ternary Bonsai 27B at 1.71 effective bits per weight and 5.9 GB, and 1-bit Bonsai 27B at 1.125 bits per weight and 3.9 GB, which the company says fits the roughly 6 GB app memory budget of an iPhone 17 Pro. Both carry a 262K-token context, support speculative decoding, and pair the language network with a 4-bit vision tower. On a 15-benchmark suite in thinking mode, the ternary variant scores 80.5 overall versus 85.0 for the full-precision baseline (95% retention) and the 1-bit variant 76.1 (90%). Weights are on Hugging Face under Apache 2.0, running via MLX on Apple devices and CUDA on NVIDIA GPUs, with a free limited-time developer preview API and quoted throughput up to 163 tok/s in 1-bit on an RTX 5090.

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.