Article may be outdated

This article is 46 days old. Some details may have changed since publication.

Hacker News·4 min read·hard

Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

P
pich
Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP
✦AI Summary

A technical experiment demonstrates how to optimize the Qwen3.8-27B AI model for local inference on consumer-grade hardware. The author achieves significant throughput gains by balancing model precision, speculative decoding, and memory management.

Why it matters

Efficient local inference allows for powerful AI capabilities on smaller, accessible hardware, reducing reliance on expensive cloud-based APIs.

✦Dive DeeperCreate a free account to unlock

I gave Qwen3.8's MTP drafter another 69.2 MiB of precision. Throughput fell from 50.44 to 37.02 tokens per second. That result sums up the whole experiment: the best local inference setup is rarely made from the individually "best" parts.

I wanted a dense 27B model, its full 262,144-token context, multimodal input, maximum useful quality, and speculative decoding on an NVIDIA RTX PRO 4000 Blackwell SFF with 24 GB of VRAM. The server also had to survive real agent work after printing model loaded . The experiment followed a hunch I had written about earlier : careful operation may matter as much as moving to a larger model.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyscience
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in