ds4-server
Local OpenAI/Anthropic-compatible HTTP server for DwarfStar LLM inference
TLDR
SYNOPSIS
ds4-server [options]
DESCRIPTION
ds4-server is the HTTP API front end of DwarfStar (project ds4 by Salvatore Sanfilippo / antirez and contributors). It loads a project-specific GGUF and serves OpenAI- and Anthropic-compatible endpoints so local tools, IDEs, and coding agents can talk to a specialized DeepSeek V4 (and related) inference engine without calling a cloud API.Each client connection is handled by a blocking request thread; inference itself is serialized onto a single worker that owns the live session and KV state. That design keeps session reuse, optional disk KV checkpoints, and graph execution in one place. Endpoints include /v1/chat/completions, /v1/responses, /v1/completions, and /v1/messages. Model name aliases such as deepseek-v4-flash and deepseek-v4-pro both refer to the GGUF currently loaded.Unlike generic GGUF runners, DwarfStar targets a narrow set of carefully prepared weights (DeepSeek V4 Flash/PRO and experimental forks for other MoE models). Use the project's download scripts and published GGUFs; arbitrary community GGUF files are not expected to work. Backends include Metal (primary on Apple Silicon), CUDA (including DGX Spark), and ROCm (e.g. Strix Halo). --ssd-streaming keeps models runnable when they exceed RAM by paging routed experts from disk.
PARAMETERS
-m, --model FILE
Path to the GGUF model. Default: ds4flash.gguf.--metal | --cuda | --rocm | --cpu
Select the inference backend explicitly. Metal is primary on macOS; CUDA/ROCm on Linux where built.-c, --ctx N
Allocated context length in tokens.-n, --tokens N
Default maximum output tokens when the client does not set a limit.--host HOST
Bind address. Default: 127.0.0.1.--port N
Bind port. Default: 8000.--cors
Add Access-Control-Allow-* headers for browser JavaScript clients.--power N
GPU duty-cycle target from 1 to 100. Default: 100.--ssd-streaming
Opt into SSD-backed model streaming instead of full RAM residency.--ssd-streaming-cache-experts N|NGB
Size of the routed-expert cache (expert count or GiB, e.g. 32GB).--think / --think-max / --nothink
Default thinking mode for chat-style requests (server also maps OpenAI/Anthropic effort fields).--kv-disk-dir DIR
Enable on-disk KV checkpoints in DIR.--kv-disk-space-mb N
Disk budget in megabytes when KV disk is enabled. Default: 4096.--trace FILE
Write prompts, cache decisions, outputs, and tool calls to a trace file.-t, --threads N
CPU helper threads for host-side work.
CAVEATS
Software is explicitly beta. Only DwarfStar-prepared GGUFs are supported. Default bind is localhost; expose to a network only if you understand there is no built-in auth. On-disk KV and large contexts need substantial free disk and RAM. CPU-only builds are mainly for diagnostics; production use expects Metal, CUDA, or ROCm. The interactive CLI binary is also named ds4, which collides with unrelated DualShock 4 tools of the same name on Linux.
HISTORY
DwarfStar (ds4) was created by Salvatore Sanfilippo (antirez) as a small, self-contained local inference stack optimized for large open-weight MoE models that barely fit (or do not fit) in consumer RAM. ds4-server provides the HTTP API half of that stack so coding agents and OpenAI-compatible clients can use the same engine as the interactive ds4 CLI.