Article may be outdated

This article is 3 days old. Some details may have changed since publication.

Hacker News·3 min read·medium

Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s

C
carloslfu
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
AI Summary

A developer has released a tool called slotstream that allows users to run the massive 104GB Qwen3.8-Flash-Next AI model on Apple Silicon Macs with limited RAM. The tool uses SSD streaming to bypass memory constraints, though it requires significant free disk space.

Why it matters

This demonstrates a significant breakthrough in local AI inference, enabling high-end model performance on consumer-grade hardware.

Dive DeeperCreate a free account to unlock

Run Qwen3.8-Flash-Next on a Mac that cannot hold it. The model is 104 GB at 4-bit; slotstream streams it from SSD and runs it in whatever memory you give it, down to an 8.1 GB planned floor. One Swift binary with the commonly used Ollama and OpenAI chat/generate endpoints.

Disk is the gate that bites first. You need ~110 GB free, so a 512 GB Mac is the realistic minimum however much memory it has. The weights are a one-time 104 GB download: well under an hour on a fast connection, several hours on a slow one (table below ).

Only the 48 GB row is measured on real hardware; the rest come from the same measured curve, and smaller Macs also have slower SSDs. Run slotstream doctor to see what your machine would get, and whether you have the disk for the weights, before downloading anything.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in