UNLTD Inference Review: Disk-First CPU-First LLM Inference in Rust

UNLTD Inference Review: Disk-First CPU-First LLM Inference in Rust

UNLTD Inference is a disk-first, CPU-first LLM runtime in Rust that runs GGUF models larger than its configured memory budget by memory-mapping weights instead of copying them to the heap. Version 1.0.0 validates one model, Ornith 1.0 9B Q4_K_M (5.63 GB), on Windows x86-64 with scalar kernels, and publishes ~73 s prefill and ~14.3 s per token. That is the entire honest summary, and everything interesting about the project lives in the gap between the headline claim and the fine print. The headline is “a 5.63 GB model on a 3 GB budget.” The fine print is that the budget governs about 113 MB of runtime-controlled memory while peak working set still reaches ~5.24 GB, and that “disk-first” here means memory-mapped files plus OS demand paging, the same mechanism llama.cpp has shipped by default for years. ...

October 1, 2026 · 19 min · baeseokjae