From the creator of Redis; run LLM locally with ds4

DwarfStar 4 (ds4) is a new C-based inference engine designed to run large language models locally on high-memory hardware. It supports specific model families like DeepSeek V4 and Qwen3.8, offering CLI, API, and agent-based interfaces.
Why it matters
This tool enables developers to run sophisticated, large-scale AI models on local hardware rather than relying on expensive cloud-based infrastructure.
DwarfStar 4 is a narrow C inference engine for high-memory Mac, CUDA and ROCm machines. It supports DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, with text and vision models, local APIs, a CLI and a native agent in one stack.
SUPPORTED: DEEPSEEK V4 / V4.1 + GLM 5.x + QWEN3.8 · MIT LICENSE · C / METAL / CUDA / ROCM · QWEN ON 64GB
DeepSeek V4 Flash is a large mixture-of-experts model. The usual path is remote serving; ds4 starts from the opposite constraint.
Asymmetric quantization targets the routed experts while preserving critical paths. The model becomes practical on high-memory machines.
The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.
Not a generic GGUF runner. ds4 follows a small, opportunistic set of model families and validates each supported layout end to end.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in