Ollama
One-command local LLM runtime.
The "A1111 for LLMs" — multi-loader local chat UI.
Runs locally · High-end GPU (16–24 GB)
EXL2 on 24 GB cards is the sweet spot for 70B Q3.
Source is public — you can audit it, fork it, and you'll never lose access to your workflows if Text Generation WebUI the company changes direction.
Fits on entry-level cards (GTX 1660, RTX 3050, RTX 4060). Rare for this category.
Native Metal / MPS support — runs on M-series Macs without CUDA gymnastics.
Power-user score 86/100 — consistently rated highly by people who use this every day, not just benchmark chasers.
This page summarizes upstream documentation, release information, and editorially reviewed catalogue fields. It is not presented as a hands-on benchmark. Verify changing requirements at the official project; report stale data through our corrections channel.
Runs CPU-only — no CUDA / driver gymnastics required.
Fits on 6 GB cards — GTX 1660 / RTX 3050 territory.
Open-source AND ships an API — easy to integrate, possible to host yourself.
Catalogue entry last updated 81 days ago — re-verification due soon.
Hand-picked from YouTube, Reddit, GitHub, and the wider web. Each link goes straight to the source — we don't intercept or rewrite anything.
Three picks across different tradeoffs — so you don't end up with three near-clones of Text Generation WebUI.
Oobabooga's gradio UI for local LLMs. Supports llama.cpp, ExLlamaV2, Transformers, and more. The go-to power-user chat front-end for hobbyists running quantized 70B models on consumer GPUs.
Free / OSS.