Hacker News·4 min read·hard

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

S
snehesht
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
✦AI Summary

A new open-source tool called Strata allows users to run the 125-billion-parameter Qwen 3.8 Flash-Next AI model on consumer-grade gaming PCs. The software requires at least 12GB of VRAM and aims to make large-scale AI accessible without the need for enterprise-grade server hardware.

Why it matters

This development democratizes access to high-performance AI, enabling local, private execution of large models on personal hardware.

✦Dive DeeperCreate a free account to unlock

English · 简体中文 · 日本語 · Deutsch · Français · Español · Português

Run a 125-billion-parameter AI model on your own gaming PC NVIDIA or AMD graphics card (12 GB or more) · Windows or Linux · free and open source

A voxel pagoda garden, 1 shot prompt running on an RTX 5070 with Strata (IQ3_S, 128K context) · full video (49 s)

Strata runs Qwen3.8-Flash-Next on a normal PC. This is a large, smart AI model that usually needs a server. It chats, writes code, reads pictures and works with your apps and coding agents. Nothing leaves your PC.

We measured it on two ordinary gaming PCs. A token is about ¾ of a word.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in