Article may be outdated

This article is 58 days old. Some details may have changed since publication.

Hacker News·4 min read·hard

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

L
leonickson
Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
✦AI Summary

A new software project called Swiftlet allows large language models like Qwen 80B to run on consumer Apple hardware by streaming model weights from storage. This enables high-parameter models to function on devices with limited RAM, such as iPhones.

Why it matters

This breakthrough demonstrates significant progress in local AI inference, potentially democratizing access to powerful models without requiring expensive cloud infrastructure.

✦Dive DeeperCreate a free account to unlock

Run 35B and 80B Qwen models on ordinary Apple devices, including iPhones.

Swiftlet is a Swift + Metal runtime for the Qwen3-Next and Qwen3.5/3.6 MoE hybrid model family. It keeps only the small dense core of a model resident in memory and streams the routed Mixture-of-Experts weights from storage on demand. The result:

The 35B also runs on an iPhone 17 in about 2.5 GB of RAM, at about 1 tok/s today. As far as we know, that is the first time a model of this class has run natively on a phone.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in