Article may be outdated

This article is 63 days old. Some details may have changed since publication.

Hacker News·4 min read·hard

Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac

G
gitpusher42
Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
✦AI Summary

A developer has released an open-source engine called TurboFieldfare that allows running a 26-billion-parameter AI model on Apple Silicon Macs with only 8 GB of RAM. It achieves this by streaming model experts from the SSD rather than loading the entire model into memory.

Why it matters

This represents a significant technical breakthrough in local AI inference, making powerful models accessible on consumer-grade hardware.

✦Dive DeeperCreate a free account to unlock

Gemma 4 26B-A4B inference in about 2 GB of RAM A custom Swift + Metal runtime for any Apple Silicon Mac, even the 8 GB ones.

Quick start · Local server · Benchmarks · Contribute results · How it works · Experiments · References

Memory got expensive. So I gave a 26-billion-parameter model a ~2 GB budget.

TurboFieldfare runs the instruction-tuned Gemma 4 26B-A4B without loading the entire 14.3 GB model into memory. It keeps the shared 1.35 GB core and FP16 KV cache in memory, then streams only the experts needed for each token from SSD. This is what lets the model run on Macs with 8 GB of RAM.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in