llama.cpp
The C++ inference engine powering most local LLMs.
OPEN SOURCECPU-CAPABLE
Absurdly cheap open-weights workhorse — $0.10/$0.20 per 1M.
DeepSeek V4 Flash is the budget king: $0.10/$0.20 per 1M tokens for a capable 1M-context open-weights model. It is the default model for this site's own automation pipeline — good enough for structured drafting at near-zero cost.
High-volume drafting, extraction and automation where cost rules.
Self-host quantised builds via llama.cpp / Ollama on 24GB+ GPUs.
The C++ inference engine powering most local LLMs.