A new local inference engine called DwarfStar 4, or ds4, lets users run large AI models on their own machines instead of sending requests to a distant server. The project is from the creator of Redis.
What DwarfStar 4 Does
The DwarfStar 4 engine runs on Mac systems with ample memory, along with CUDA and ROCm hardware, and it works with DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, including text and vision models, local APIs, a command-line interface and a native agent all together in one package. This inference engine operates within a narrow C framework and ships under the MIT License.
Compression sits at the center of the design. The software compresses the routed experts inside a mixture-of-experts model, all while preserving the shared paths that matter most. This is what makes the supported routed-MoE builds match their target machines.
The Three Phases
The project breaks its development path into three phases:
- Phase 1: The Giant — DeepSeek V4 Flash starts as a large mixture-of-experts model served remotely.
- Phase 2: The Collapse — Asymmetric quantization targets the routed experts while preserving critical paths, making the model practical on high-memory machines.
- Phase 3: The Dwarf Star — The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.
How It Works
DwarfStar 4 does not function as an ordinary GGUF runner. Instead, it relies on a limited selection of model families and tests every supported layout from start to finish. The approach includes keeping lengthy prefixes on SSD and restarting based on prompt hash, which means that a session can be picked up without repeating the entire initial setup.
A self-contained engine sits at the core of the architecture, paired with agent-facing interfaces, while Project GGUFs serves as the foundation, all verified against official model outputs.
Running It Yourself
The process is simple:
- Fetch the weights using the
./download_model.shscript. - Build for your backend using
make. - Talk to the model through the CLI or start the server.
A benchmark table offers the numbers behind the project, and among them sits the M5 MAX 128GB reference row, which carries an 32K CTX figure alongside a generation rate of 34.4 T/S and a prefill rate of 557 T/S. The complete guide is found in Hardware and Installation.
Design Choices
The whole point of DwarfStar 4 is stated plainly and repeated across several sections: run models locally without sending requests to a distant server. The software compresses the routed experts in a mixture-of-experts model while keeping critical shared paths precise.
This section spells out the design choices: the software sticks to a limited number of model families and checks each arrangement from start to finish, rather than relying on a generic GGUF runner.
DwarfStar 4 is a working local model stack from a team with a track record in distributed systems.
Source material: “From the creator of Redis; run LLM locally with ds4,” dwarfstar.sh.
Get the Notebook.
The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

