Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
A new software project called Swiftlet allows large language models like Qwen 80B to run on consumer Apple hardware by streaming model weights from storage. This enables high-parameter models to function on devices with limited RAM, such as iPhones.
Why it matters
This breakthrough demonstrates significant progress in local AI inference, potentially democratizing access to powerful models without requiring expensive cloud infrastructure.
Run 35B and 80B Qwen models on ordinary Apple devices, including iPhones.
Swiftlet is a Swift + Metal runtime for the Qwen3-Next and Qwen3.5/3.6 MoE hybrid model family. It keeps only the small dense core of a model resident in memory and streams the routed Mixture-of-Experts weights from storage on demand. The result:
The 35B also runs on an iPhone 17 in about 2.5 GB of RAM, at about 1 tok/s today. As far as we know, that is the first time a model of this class has run natively on a phone.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in