PrismML released Bonsai 27B on July 14, 2026, a multimodal model compressed from Qwen3.6 27B and shipped in two variants: Ternary Bonsai 27B at 1.71 effective bits per weight and 5.9 GB, and 1-bit Bonsai 27B at 1.125 bits per weight and 3.9 GB, which the company says fits the roughly 6 GB app memory budget of an iPhone 17 Pro. Both carry a 262K-token context, support speculative decoding, and pair the language network with a 4-bit vision tower. On a 15-benchmark suite in thinking mode, the ternary variant scores 80.5 overall versus 85.0 for the full-precision baseline (95% retention) and the 1-bit variant 76.1 (90%). Weights are on Hugging Face under Apache 2.0, running via MLX on Apple devices and CUDA on NVIDIA GPUs, with a free limited-time developer preview API and quoted throughput up to 163 tok/s in 1-bit on an RTX 5090.
Sources