Running Gemma 4 26B at 5 tokens/SEC on a 13-year-old Xeon with no GPU

A tech enthusiast successfully ran a modern 26-billion-parameter AI model on 13-year-old server hardware without a GPU. The project highlights the engineering challenges of optimizing software for older CPU architectures lacking modern instruction sets.
Why it matters
This demonstrates that high-level AI capabilities can be democratized through clever software optimization, reducing the reliance on expensive, modern hardware.
There’s a server in my basement that has no business running a modern language model. It’s a repurposed HP StoreVirtual storage box, roughly thirteen years old, two Ivy Bridge Xeons, no GPU. It was built to hold disks, not do math. As of this week it runs Google’s Gemma 4 , a 26-billion-parameter open-weights mixture-of-experts model, at about five tokens per second. Reading speed.
The article is a personal technical account focused on engineering experimentation rather than ideological bias.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in