waste
local LLM inference that streams model experts from disk
TLDR
SYNOPSIS
waste command [options] [args...]
DESCRIPTION
waste (Weight-Aware Streaming Tensor Engine) is a dependency-free C inference engine and CLI for running large mixture-of-experts language models when the full weight set does not fit in RAM. It keeps a resident "trunk" in memory, streams activated experts from a converted .waste container on NVMe, and uses remaining RAM as a bounded expert cache.The flagship proof point is open-weight Kimi K3 (~2.78T parameters, ~982 GiB container) running on consumer hardware at roughly half a token per second with enough RAM and internal NVMe. Smaller models in the same format (e.g. Kimi-Linear 48B) run much faster with far lower RAM floors.Convert published safetensors with the repo's Python tools once; at runtime only the waste binary (and libc/pthreads) is required. An optional OpenAI-compatible HTTP server lives under `serve/` and uses the same public C API via ctypes.
PARAMETERS
run CONTAINER [PROMPT]
Generate a completion from a prompt (or stdin). Common flags: -n count (token cap), --budget SIZE, --image FILE, --verify.chat CONTAINER
Multi-turn interactive session; conversation state is kept in the process.eval CONTAINER PROMPT
Show next-token scores/distribution. Supports --top-k, --json, --image.plan CONTAINER
Report what RAM budget fits on this machine and what the engine would pick.info CONTAINER
Print container/model metadata.bench CONTAINER
Run a built-in performance check.tokenize CONTAINER
Tokenize or inspect prompt layout (see --help for current subflags).--budget SIZE
Hard RAM ceiling (e.g. `46G`). If omitted, the engine picks a budget under about 7/8 of physical RAM, never below the model floor.--verify
Check each expert record CRC on cache miss (slower; useful after copy/download). Also WASTE_VERIFY=1.--json
Machine-readable output for eval/plan/info/bench and similar commands.--help
List all commands and flags.
CAVEATS
K3-class models need tens of GB of RAM (about 29 GB minimum open, ~64 GB for useful throughput) and a ~1 TB container on fast internal NVMe—USB enclosures are far too slow. Build needs a C11 compiler and make; conversion needs Python/torch. Expert CRC checks are off by default. Chat templates are fully filled for models the converter knows (K3 today); other containers may run in raw prompt mode. Not a drop-in replacement for general multi-model runtimes like ollama.
HISTORY
WASTE is developed by SQLite Cloud, Inc. (sqliteai) and released under Apache 2.0. It targets disk-streaming inference for frontier-scale MoE models on single machines rather than multi-GPU servers.