Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
A developer has released an open-source engine called TurboFieldfare that allows running a 26-billion-parameter AI model on Apple Silicon Macs with only 8 GB of RAM. It achieves this by streaming model experts from the SSD rather than loading the entire model into memory.
Why it matters
This represents a significant technical breakthrough in local AI inference, making powerful models accessible on consumer-grade hardware.
Gemma 4 26B-A4B inference in about 2 GB of RAM A custom Swift + Metal runtime for any Apple Silicon Mac, even the 8 GB ones.
Quick start · Local server · Benchmarks · Contribute results · How it works · Experiments · References
Memory got expensive. So I gave a 26-billion-parameter model a ~2 GB budget.
TurboFieldfare runs the instruction-tuned Gemma 4 26B-A4B without loading the entire 14.3 GB model into memory. It keeps the shared 1.35 GB core and FP16 KV cache in memory, then streams only the experts needed for each token from SSD. This is what lets the model run on Macs with 8 GB of RAM.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in