Midterms 2026See who we think should earn your vote, based on our standardsThe guide →
WRITTEN IN PLAIN AMERICAN ENGLISH.
CLAY TRIBUNE.
Advertisement

From the Creator of Redis: Run LLM Locally With DwarfStar 4

A narrow engine runs large models locally, compressing routed experts while keeping shared paths precise.

By mitch·3 min read
A glowing motherboard glows with tiny stars of light in a dark data center corridor.

A new local inference engine called DwarfStar 4, or ds4, lets users run large AI models on their own machines instead of sending requests to a distant server. The project is from the creator of Redis.

What DwarfStar 4 Does

The DwarfStar 4 engine runs on Mac systems with ample memory, along with CUDA and ROCm hardware, and it works with DeepSeek V4 and V4.1 Flash, GLM 5.x and Qwen3.8 Flash Next, including text and vision models, local APIs, a command-line interface and a native agent all together in one package. This inference engine operates within a narrow C framework and ships under the MIT License.

Compression sits at the center of the design. The software compresses the routed experts inside a mixture-of-experts model, all while preserving the shared paths that matter most. This is what makes the supported routed-MoE builds match their target machines.

Advertisement

The Three Phases

The project breaks its development path into three phases:

  1. Phase 1: The Giant — DeepSeek V4 Flash starts as a large mixture-of-experts model served remotely.
  2. Phase 2: The Collapse — Asymmetric quantization targets the routed experts while preserving critical paths, making the model practical on high-memory machines.
  3. Phase 3: The Dwarf Star — The local engine exposes a CLI, HTTP APIs and a native agent, all sharing the same model state and cache.

How It Works

DwarfStar 4 does not function as an ordinary GGUF runner. Instead, it relies on a limited selection of model families and tests every supported layout from start to finish. The approach includes keeping lengthy prefixes on SSD and restarting based on prompt hash, which means that a session can be picked up without repeating the entire initial setup.

A self-contained engine sits at the core of the architecture, paired with agent-facing interfaces, while Project GGUFs serves as the foundation, all verified against official model outputs.

Running It Yourself

The process is simple:

  1. Fetch the weights using the ./download_model.sh script.
  2. Build for your backend using make.
  3. Talk to the model through the CLI or start the server.

A benchmark table offers the numbers behind the project, and among them sits the M5 MAX 128GB reference row, which carries an 32K CTX figure alongside a generation rate of 34.4 T/S and a prefill rate of 557 T/S. The complete guide is found in Hardware and Installation.

Design Choices

The whole point of DwarfStar 4 is stated plainly and repeated across several sections: run models locally without sending requests to a distant server. The software compresses the routed experts in a mixture-of-experts model while keeping critical shared paths precise.

This section spells out the design choices: the software sticks to a limited number of model families and checks each arrangement from start to finish, rather than relying on a generic GGUF runner.

DwarfStar 4 is a working local model stack from a team with a track record in distributed systems.

Source material: “From the creator of Redis; run LLM locally with ds4,” dwarfstar.sh.

The Notebook

Get the Notebook.

The day's best stories and every fresh verdict, in plain English, in your inbox by seven. One email a day, no more.

We send one note to confirm. Every issue has a one-click way out.

Advertisement

Leave a Reply

Your email address will not be published. Required fields are marked *

As an Amazon Associate, Clay Tribune earns from qualifying purchases.