llama.cpp
The C++ inference engine powering most local LLMs.
Running your coding assistant locally keeps your codebase off third-party servers and removes rate-limit anxiety. The right pick is more about the runner than the model — they're all interchangeable at the API layer.
The catalogue picks below are a shortlist, not proof that every default configuration fits. Open each datasheet and verify the exact model or extension you intend to use.
See the assumptions, official sources, and memory trade-offs behind this workflow.
The C++ inference engine powering most local LLMs.
One-command local LLM runtime.
RAG-first local LLM workspace with workspaces and agents.
Desktop GUI for running local LLMs.
Beautifully designed chat UI with plugins and image generation.
Self-hosted ChatGPT-style frontend for Ollama / OpenAI.