Hacker News·5 min read·hard

From the creator of Redis; run LLM locally with ds4

F
fibo
From the creator of Redis; run LLM locally with ds4
✦AI Summary

DwarfStar 4 (ds4) is a new C-based inference engine designed to run large language models locally on high-memory hardware. It supports specific model families like DeepSeek V4 and Qwen3.8, offering CLI, API, and agent-based interfaces.

Why it matters

This tool enables developers to run sophisticated, large-scale AI models on local hardware rather than relying on expensive cloud-based infrastructure.

✦Dive DeeperCreate a free account to unlock

DwarfStar 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.

SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE · C / METAL / CUDA / ROCM · QWEN ON 64GB

DeepSeek V4 Flash is a large mixture-of-experts model. The usual path is remote serving; ds4 starts from the opposite constraint.

Asymmetric quantization targets the routed experts while preserving critical paths. The model becomes practical on high-memory machines.

The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.

Not a generic GGUF runner. ds4 follows a small, opportunistic set of model families and validates each supported layout end to end.

Continue reading on Headlinne

Create a free account to read the full article.

Read full article →
technologyai
✦

Get smarter about the news

Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.

Create free account

Already have an account? Sign in