Key takeaways

  • If AI lab PrismML isn’t on your radar yet, it should be — not because it’s raised gobs of money (it hasn’t yet, just a $22.25 million seed…
  • 25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it’s developing.
  • It’s a 9x to 10x reduction in memory versus the original.

What happened

25 million seed round), but because of the technical minds involved and the potentially industry-changing tech it’s developing. PrismML is betting that capable, high-performing, reasoning large language models don’t, in fact, have to be large. It is making reasoning models so small they can fit on PCs and smartphones. 9 GB. That’s small enough to fit on a PC and, possibly, a high-end smartphone.

” Stoica tells us that he’s excited for this tech because it’s making it possible for advanced models to run on users’ devices. “You are going to have intelligence at your fingertips, and it’s going to be free because it’s going to run on the device you already bought. ” When you purchase through links in our articles, we may earn a small commission. This doesn’t affect our editorial independence.

Why it matters

It’s a 9x to 10x reduction in memory versus the original. PrismML was founded by a group of Caltech researchers and is led by Hassibi, a Caltech professor and an expert in compression technologies. The startup also counts Ion Stoica as an adviser.

Stoica is a co-founder of Databricks (and other companies) and the director of Berkeley’s famed Sky Computing Lab, which has birthed many technologies and startups, from Letta to SGLang. PrismML is also backed by investors Khosla Ventures, Cerberus Capital, and Caltech. This startup is certainly not the only company working on LLM compression tech. Multiverse Computing, founded by a well-known professor from Spain’s Donostia International Physics Center, is another.

What to watch

Whether it could ever get to 100% benchmark performance parity is a question that remains to be seen. Compression will likely always have some impact, Hassibi says. Still, perfect benchmark parity is fairly academic anyway. LLMs are not so accurate in their uncompressed form, and benchmarks not so perfectly reflective of actual tasks, that a 2% degradation would likely meaningfully affect how a model performs in actual use.