Unleashing On-Device AI: The Future is Small and Smart
The Vision: Powerful AI in Your Pocket
PrismML, an AI research company, is gaining recognition for its groundbreaking work, not because of massive funding rounds (having secured a modest $22.25 million seed round), but due to the caliber of its technical team and the potentially transformative technology it's developing.
Challenging the Notion of Scale in LLMs
At the core of PrismML's philosophy is the belief that highly capable, high-performing, and intelligent large language models do not inherently need to be massive in size. They are proving that advanced reasoning models can be shrunk to fit within the confines of personal computers and even smartphones.
Introducing Bonsai 2 27B: A Leap in Compression
PrismML recently unveiled Bonsai 2 27B, the newest addition to its model family. This model effectively compresses Qwen3.8 27B, a widely utilized open-source model from Alibaba, to a mere 5.9 GB. This significant reduction in memory, approximately 9 to 10 times less than the original, makes it suitable for PCs and potentially high-end smartphones. Reports suggest that PrismML might be in discussions with major tech companies like Apple regarding this innovative compression technology.
The Minds Behind PrismML's Innovation
The company was founded by a group of distinguished Caltech researchers, with Caltech professor Babak Hassibi, an expert in compression techniques, at the helm. PrismML also benefits from the advisory role of Ion Stoica, a co-founder of Databricks and director of Berkeley's renowned Sky Computing Lab, a hub for technological innovation and startup creation. PrismML's development is further supported by investments from prominent firms such as Khosla Ventures and Cerberus Capital, alongside Caltech itself.
Unrivaled Performance Retention in Compression
While other companies are also exploring LLM compression, Hassibi asserts that PrismML's methodology is unique due to its ability to maintain virtually identical performance compared to the original, uncompressed models. Bonsai 2, for instance, achieves 98% of Qwen's aggregated benchmark scores, an improvement from its predecessor, the first Bonsai, which hit 95%. The initial Bonsai model has already seen over 11 million downloads, with PrismML's even smaller models accumulating an additional 2.6 million downloads.
The Path to Perfect Benchmark Parity and Future Goals
This consistent improvement in compression performance across releases indicates PrismML's progress. While achieving 100% benchmark performance parity remains a future challenge, Hassibi acknowledges that some impact from compression is inherent. Nevertheless, such minor degradations are largely inconsequential given the inherent inaccuracies of uncompressed LLMs and the practical limitations of benchmarks in reflecting real-world performance. The accompanying software that houses the model, known as the harness, also plays a crucial role in overall accuracy.
The 'Ternary' Weight Approach to Compact AI
PrismML's success stems from its method of shrinking a model's 'weights'—the stored information acquired during training. Traditional weights typically require 16 bits, but PrismML's innovative 'ternary' weights simplify this to just three values: +1, -1, or 0. This drastically reduces the storage space needed for each weight, leading to significantly smaller models. Further details on this compression technique can be found on their Hugging Face page.
Expanding Horizons: Compression for Larger Models
The startup's next ambition is to apply this powerful compression technique to even larger models. Hassibi anticipates releasing models in the several-hundred-billion-parameter range in the coming months, believing that preserving intelligence becomes easier with increased model size. He suggests a general trend where achieving 100% performance retention is more attainable with larger models.
The Promise of Accessible and Private Intelligence
Stoica expresses great enthusiasm for this technology, highlighting its potential to enable advanced AI models to run directly on user devices. This development promises to make intelligent capabilities freely accessible to everyone, leveraging existing hardware, while also ensuring enhanced privacy by keeping data on-device rather than relying on cloud-based processin
