Research

Apple in Early Talks with PrismML on Running Compressed LLMs Locally on iPhone

Apple is in preliminary discussions with PrismML, a Caltech spin‑off backed by Khosla Ventures, which says it has compressed large language models so they can run on iPhone 15 and newer devices.

Apple in Early Talks with PrismML on Running Compressed LLMs Locally on iPhone

Apple is in preliminary discussions with PrismML, a California startup that says it has compressed large language models sufficiently to run them directly on iPhones, PrismML CEO Babak Haszibi told CNBC. PrismML is a Caltech spin‑off backed by Khosla Ventures.

What the compression claims mean

On Tuesday PrismML published compressed versions of Alibaba’s open‑source Qwen model. The startup says it reduced an original model of roughly 54 gigabytes to under 4 gigabytes, enabling all 27 billion parameters to run on an iPhone 15 or newer device.

PrismML’s method dramatically simplifies internal data storage, reducing individual values from 16 bits to one or three possible values. According to the company’s figures, the compressed models require 10–15 times less memory, respond 6–8 times faster, and consume 3–6 times less energy than the conventional versions. The company also acknowledges a trade‑off: overall performance typically declines by a few percentage points, with factual recall — and first impacts on reasoning, mathematical ability and coding — often showing the largest degradation.

Why Apple would be interested and timing

The announcement came one day after Apple released the iOS 27 public beta, which expands access to the redesigned Siri for more iPhone users. Apple’s stated aim is to make Siri more competitive with offerings from OpenAI and Anthropic while keeping as much computation and AI data processing on‑device as possible for privacy reasons.

On‑device models would reduce network latency, lower cloud infrastructure costs and strengthen Apple’s privacy arguments; some features could also work without an internet connection. Haszibi said Apple and other technology companies are already testing PrismML’s solutions for speed, energy efficiency and general performance, but described the talks as early stage and said they are “moving along nicely.”

IP, funding and next steps

PrismML’s technology originates from a research group at Caltech led by Haszibi, and Caltech has granted PrismML exclusive usage rights to the related patents. The startup closed a $16.25 million financing round in March. PrismML says the next models in its roadmap include Google’s open‑source Gemma, followed by work on larger models.

Analyst caution and technical risks

Analysts urge caution: PrismML’s claims need to be validated beyond laboratory demos. Counterpoint Research experts say performance on complex instructions, battery usage under multitasking conditions, and reliability across millions of queries will be decisive. An IDC specialist warned that energy consumption may be the biggest challenge since models running frequently or continuously in the background can drain batteries quickly even if they need less memory.

A D.A. Davidson analyst noted that shrinking model sizes may not eliminate demand for chips but rather shift that demand from data centers to mobile devices, where chip utilization can remain low and overall efficiency may be worse.

Conclusion

PrismML’s compression approach could enable more powerful AI features to run directly on iPhones, reducing latency and improving privacy. However, whether the technology is practical at scale depends on real‑world testing of battery impact, multitasking performance and long‑term reliability.