Qwen3.8-27B at 256K on a 24GB RTX PRO 4000 SFF (432 GB/s): 50 tok/s with MTP

A technical experiment demonstrates how to optimize the Qwen3.8-27B AI model for local inference on consumer-grade hardware. The author achieves significant throughput gains by balancing model precision, speculative decoding, and memory management.
Why it matters
Efficient local inference allows for powerful AI capabilities on smaller, accessible hardware, reducing reliance on expensive cloud-based APIs.
I gave Qwen3.8's MTP drafter another 69.2 MiB of precision. Throughput fell from 50.44 to 37.02 tokens per second. That result sums up the whole experiment: the best local inference setup is rarely made from the individually "best" parts.
I wanted a dense 27B model, its full 262,144-token context, multimodal input, maximum useful quality, and speculative decoding on an NVIDIA RTX PRO 4000 Blackwell SFF with 24 GB of VRAM. The server also had to survive real agent work after printing model loaded . The experiment followed a hunch I had written about earlier : careful operation may matter as much as moving to a larger model.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in