PrismML Shrinks a 27B Reasoning Model to 5.9 GB for Local AI

PrismML released Bonsai 2 27B, a compressed reasoning model designed for PCs and phones, strengthening the case for private, on-device AI.

AI News

PrismML has released Bonsai 2 27B, a compressed reasoning model that the startup says can deliver nearly the performance of a much larger model while fitting into 5.9 GB of storage. The release is the clearest test yet of PrismML’s effort to move capable language models from cloud servers onto PCs and, potentially, high-end smartphones.

The company’s approach could matter to AI builders and product teams because local inference can reduce cloud costs, improve response latency, and keep sensitive data on a user’s device. But the strongest performance and adoption figures available so far come from PrismML itself, and the startup has not demonstrated that benchmark results will translate equally across real-world applications.

A smaller model with a larger ambition

Bonsai 2 27B compresses Qwen3.8 27B, an open-source model from Alibaba, into a file that is roughly 9 to 10 times smaller than the original according to TechCrunch’s reporting. At 5.9 GB, the model should fit comfortably on many personal computers and may run on some high-end phones, although practical performance will depend on memory, processor capabilities, quantization, and the software used to run it.

PrismML was founded by Caltech researchers and is led by Babak Hassibi, a Caltech professor known for work in compression. Ion Stoica, a Databricks co-founder and director of the University of California, Berkeley’s Sky Computing Lab, is an adviser. The startup has raised a $22.25 million seed round from Khosla Ventures, Cerberus Capital, and Caltech.

The company is not alone in pursuing LLM compression. Multiverse Computing is also developing methods to reduce model size. PrismML’s distinction, according to Hassibi, is that its models retain nearly all of the source model’s measured capability rather than trading substantial accuracy for a smaller footprint.

How PrismML says the compression works

PrismML’s method uses what it calls ternary weights. Model weights normally store learned information using comparatively large numerical representations. The company reduces those values to three possible states: positive one, negative one, or zero. That limits the amount of information required to represent each weight and sharply reduces memory demands.

The resulting technique, known broadly as LLM compression, is aimed at changing where inference can happen. Instead of sending every prompt to a hosted service, an application could execute at least some tasks locally. That could support offline functionality, faster interactions, and greater control over data flows.

PrismML’s first Bonsai model, released a few months before Bonsai 2 27B, reportedly matched 95% of the source model’s aggregate benchmark scores. The new release is claimed to reach 98%. Those figures indicate progress between releases, but they are not the same as independent validation across a broad set of tasks, devices, and evaluation methods.

Evidence, benchmarks, and unresolved limits

The 98% result is a vendor-reported benchmark claim attributed to PrismML. It suggests a small gap against the original Qwen model on the company’s chosen aggregate evaluation, but it does not establish that Bonsai 2 27B will behave identically in production. Benchmarks can miss failures involving long context, tool use, domain-specific knowledge, multilingual prompts, or the orchestration layer around a model.

Hassibi acknowledged that compression will probably always introduce some performance impact. He also argued that a 2% benchmark difference may not have much practical significance because uncompressed language models are imperfect and benchmark scores do not precisely predict user outcomes. That is a reasonable engineering argument, but it remains a claim that product teams will need to test against their own workloads.

PrismML also says its original Bonsai model has been downloaded more than 11 million times, while its smaller models have received another 2.6 million downloads. Those numbers are company-reported adoption signals, not proof of sustained production use, active users, or commercial deployments. Download totals can include experiments, automated activity, and repeated pulls.

Why local inference matters to builders

If the model performs reliably on consumer hardware, Bonsai 2 27B could expand the design space for on-device AI. Developers could build assistants that continue working without a network connection, process private documents locally, or use a cloud model only when a task exceeds local capability. For enterprise buyers, that architecture could reduce exposure of confidential prompts and documents to external inference providers.

The trade-offs are equally important. A model that fits in storage may still be demanding to run at useful speed. Memory bandwidth, thermal limits, battery consumption, operating-system support, and the availability of optimized runtimes will determine whether a local model feels responsive. Application teams will also have to evaluate whether compressed reasoning is dependable enough for workflows such as coding assistance, document analysis, customer support, or AI agents.

PrismML’s longer-term plan is to compress models with several hundred billion parameters. Hassibi told TechCrunch that larger models may offer more room for compression without losing as much intelligence. If that claim holds, the company could eventually target local systems that are substantially more capable than today’s small models. It is still a forward-looking goal, not a demonstrated product capability.

What to watch next

The next important signal will be independent testing of Bonsai 2 27B on consumer PCs and smartphones, including speed, memory use, battery impact, and accuracy on practical tasks. Developers should also watch for supported runtimes, licensing details, model availability, and reproducible evaluation data.

PrismML says it hopes to release models in the several-hundred-billion-parameter range within the next couple of months. Whether those releases arrive on schedule, preserve the claimed quality, and run on commercially relevant hardware will show whether the company’s approach scales beyond a compelling compression demonstration.

Reports that PrismML may be in talks with Apple are unconfirmed; Hassibi declined to comment. Any formal device or platform partnership would provide a clearer signal about commercial deployment, but no such agreement has been established by the available evidence.

Creati.ai perspective

PrismML’s release is significant less because a 5.9 GB model automatically replaces cloud AI and more because it makes the local-versus-cloud decision more flexible. A compressed model that is slightly weaker but private, cheap to operate, and available offline may be more useful than a larger model for many narrow product workflows.

The key question is not whether Bonsai 2 27B matches 98% of a benchmark. It is whether builders can depend on it under real constraints: limited memory, changing prompts, tool calls, safety policies, and messy enterprise data. If PrismML can show that its compression survives those tests—and scales to larger models—it could make local AI a practical deployment tier rather than a specialist optimization.

Ads