Article may be outdated

This article is 38 days old. Some details may have changed since publication.

Hacker News·4 min read·hard

Running Gemma 4 26B at 5 tokens/SEC on a 13-year-old Xeon with no GPU

N
neomindryan
Running Gemma 4 26B at 5 tokens/SEC on a 13-year-old Xeon with no GPU
AI Summary

A tech enthusiast successfully ran a modern 26-billion-parameter AI model on 13-year-old server hardware without a GPU. The project highlights the engineering challenges of optimizing software for older CPU architectures lacking modern instruction sets.

Why it matters

This demonstrates that high-level AI capabilities can be democratized through clever software optimization, reducing the reliance on expensive, modern hardware.

Dive DeeperCreate a free account to unlock

There’s a server in my basement that has no business running a modern language model. It’s a repurposed HP StoreVirtual storage box, roughly thirteen years old, two Ivy Bridge Xeons, no GPU. It was built to hold disks, not do math. As of this week it runs Google’s Gemma 4 , a 26-billion-parameter open-weights mixture-of-experts model, at about five tokens per second. Reading speed.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technology
Political Bias
Center
LeftLean LCenterLean RRight
Confidence: 80%

The article is a personal technical account focused on engineering experimentation rather than ideological bias.

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in