Skip to content
[ STACK BUILDER · CONFIGURATOR ]

Build your local AI stack — matched to your card.

Tell us your hardware budget and what you want to make. We pick a consistent, working stack from our catalogue — workflow engine, model, runner, training tools. Empty slots tell you why nothing fits, instead of pretending.

CONFIGURATOR //LIVE
01 / What do you want to do?The job your stack needs to handle.
02 / What hardware do you have?We filter out anything that won't fit your card's minimum VRAM.
03 / PlatformApple Silicon only shows tools with native Metal / MPS support.
[ YOUR STACK ]2 / 2 SLOTS FILLED

Run local LLMs on 16 GB VRAM

Run language models on your own hardware — Ollama, llama.cpp, LM Studio. Cards in this tier: RTX 4070 Ti SUPER, RTX 4080, M-series 16 GB.

LLM RUNNER🖥️ Runtime

Ollama

One-command local LLM runtime.

CPU-capable · no GPU floor

Role: The runtime that loads quantised LLMs and serves them locally.

OPEN SOURCECPU-CAPABLE

ALT // Open WebUI self-hosted chatgpt-style frontend for ollama / openai.

ORCHESTRATION / SERVING🛰️ Serving

OpenAI Whisper

The reference open-source speech-to-text model.

Role: For when you outgrow a single GPU and need to serve workloads.

OPEN SOURCE2–10 GB VRAM

ALT // CrewAI role-playing agents working as a crew.

HOW WE PICK //

Every pick is a function of three things in our catalogue: minimum VRAM (must fit your budget), power-user score (60% of the weight), and trending score (20%). We add small bonuses for open-source licensing and beginner-friendly setup. No paid placements, no “sponsored” tier — if it's not in our catalogue, it can't appear here.

READ THE FULL METHODOLOGY →