Key facts
- PrismML has developed a compression technique to make large language models (LLMs) smaller.
- The company's latest model, Bonsai 2 27B, reduces a 27 billion parameter model to 5.9 GB.
- This size allows the model to fit on PCs and potentially high-end smartphones.
- PrismML's compression method uses 'ternary' weights, simplifying values to +1, -1, or 0.
- Bonsai 2 27B matches 98% of Qwen's aggregate benchmark scores.
- PrismML was founded by Caltech researchers and is led by professor Babak Hassibi.
AI startup PrismML is developing technology to significantly reduce the size of large language models (LLMs) without compromising performance, aiming to make advanced AI more accessible. The company announced its latest model, Bonsai 2 27B, which compresses Alibaba's Qwen3.8 27B model, a widely used open-source LLM, down to 5.9 GB. This reduction is a 9x to 10x decrease in memory, making the model small enough to run on personal computers and potentially high-end smartphones.
Founded by Caltech researchers and led by professor Babak Hassibi, an expert in compression technologies, PrismML has secured $22.25 million in seed funding from investors including Khosla Ventures and Cerberus Capital. The startup's unique compression technique simplifies the model's 'weights'—the learned information during training—from the standard 16 bits to just three values: +1, -1, or 0. This 'ternary' weight approach drastically reduces storage requirements.
PrismML's previous model, Bonsai, released in March, achieved 95% of Qwen's benchmark scores and has been downloaded over 11 million times, with smaller PrismML models accumulating another 2.6 million downloads. The new Bonsai 2 model improves on this, matching 98% of Qwen's aggregate benchmark scores. Hassibi acknowledges that perfect 100% parity might be unattainable due to inherent impacts of compression, but argues that the slight degradation is unlikely to affect real-world performance significantly.
Looking ahead, PrismML plans to apply its compression technique to even larger models, potentially in the several-hundred-billion-parameter range, expecting that larger models will be easier to compress while retaining intelligence. Advisor Ion Stoica, a co-founder of Databricks, highlighted the potential for advanced AI to run on users' devices, offering privacy and cost benefits by avoiding cloud reliance.
