PrismML released Ternary Bonsai 2 27B, a 5.9GB ternary-weight Qwen3.8 model

PrismML released Ternary Bonsai 2 27B on September 17, 2026, a compressed version of Alibaba's Qwen3.8 27B in which every weight is one of three values, -1, 0 or +1, with FP16 group-wise scaling. That works out to 1.76 effective bits per weight and a total footprint of 5.9GB, which PrismML says is more than 9x smaller than the full-precision model. It keeps a 262K-token context window and handles text and images.

PrismML reports an aggregate benchmark score of 83.9 against 85.4 for Qwen3.8 27B, or 98.2 percent of the original's performance. It lists speeds of up to 143 tokens per second on an NVIDIA GeForce RTX 5090 and 46.8 tokens per second on an Apple M5 Max, and energy use of 0.714 mWh per token on an RTX 4090, which it says is 40 percent more efficient than full-precision 8B models. The weights are released under Apache 2.0 and run on NVIDIA GPUs via CUDA and on Mac, iPhone and iPad via MLX.

Ternary and near-1-bit models have been an active research thread, but most results were small models trained from scratch. Compressing a capable 27B open model into a footprint that fits consumer hardware, while keeping nearly all of its benchmark score, is the practical payoff that thread was chasing, and it widens what can run locally without a data center.

The 98.2 percent figure is an aggregate chosen by PrismML; aggregate scores can hide larger drops on specific tasks such as long-context retrieval or precise reasoning, and no independent evaluation was available at release. The throughput and energy numbers are vendor measurements on specific hardware.

Sources

Last verified September 21, 2026