runnburn
Memory-aware GGUF language model inference CLI
TLDR
SYNOPSIS
runNburn [options] model.gguf [prompt]runNburn chat [options] model.ggufrunNburn serve [options] model.gguf
DESCRIPTION
runNburn (package/crate name runNburn; product binary commonly runNburn) is a pre-1.0 Rust inference runtime for quantized GGUF models that are larger than available fast memory. Weights stay file-backed where possible, host residency is bounded by detected or explicit RAM budgets, and optional CUDA or Metal paths accelerate supported operators.The same product path covers one-shot generation, interactive chat, and a local OpenAI-compatible HTTP server (chat/completions, responses, models listing). Architecture-aware support includes Llama/Phi, Gemma, Qwen dense/hybrid/MoE, and other families; recognition does not guarantee every community GGUF variant is fully accelerated.CPU is the default backend. Build with Cargo feature flags (cpu, cuda, metal, experimental vulkan) depending on hardware. Without --ram-budget, the engine reserves about one quarter of physical RAM for the OS and uses the remainder as a working-set budget.
PARAMETERS
--ram-budget size
Cap engine-owned host residency (e.g. 16GiB, 32GB). Binary and decimal suffixes are accepted. Direct CLI options must appear before the GGUF path.chat / serve
Subcommands for multi-turn REPL and HTTP serving. See runNburn chat --help and runNburn serve --help for sampling, cache, and bind options.serve options commonly include --host, --port, --model-name, --response-cache-budget, and --api-key-file. Non-loopback binds require an API key.
CAVEATS
Pre-1.0; APIs and backend coverage change. --ram-budget is not an OS RSS hard limit. The OpenAI surface is partial compatibility, not a full OpenAI API. Vulkan and some mobile paths are experimental. The binary name uses mixed case (runNburn) as shipped by the project.