ElevenLabs
The benchmark commercial TTS / voice clone API.
Local AI lives and dies by VRAM. Pick your tier — we'll show only the workflows, models, and runners that actually fit in that budget. Every entry tells you the required quantization, expected speed, and what stays out of reach.
Six gigabytes is a constrained but useful learning tier. Smaller or aggressively optimized image and language-model workflows can run, but model variant, precision, resolution, context, and offload settings decide the real fit. Leave room for runtime overhead instead of sizing from checkpoint files alone.
e.g. GTX 1660 Ti · RTX 2060 · RTX 3050 / 3060 6GB
Eight gigabytes is a practical hobbyist entry tier, not a universal compatibility line. Many image workflows and compact quantized LLMs fit with sensible settings. Official ComfyUI documentation also shows a Wan 2.2 TI2V 5B path designed for 8 GB with native offload; larger video variants remain a different workload.
e.g. RTX 3060 Ti · RTX 3070 · RTX 4060 / 4060 Ti
Twelve gigabytes gives useful headroom for complex image graphs, quantized local LLMs, and selected compact video workflows. It is still a planning tier rather than a guarantee: model architecture, context, resolution, frames, precision, and offload settings can move a workload across the boundary.
e.g. RTX 3060 12GB · RTX 4070 · RTX 4070 SUPER
Sixteen gigabytes reduces compromise for multi-component image graphs and expands quantized LLM and video options. Unified memory on Apple Silicon is not directly interchangeable with discrete VRAM: the operating system and application share the same pool, and backend support differs.
e.g. RTX 4070 Ti SUPER · RTX 4080 (16 GB) · RTX 5070 Ti
Twenty-four gigabytes is a flexible consumer-workstation tier. It supports upstream-documented Flux LoRA configurations and gives heavy image or video graphs more room, but it does not guarantee that every large model, precision, context, or frame sequence stays fully on the GPU.
e.g. RTX 3090 / 3090 Ti · RTX 4090 · RTX A5000
Forty-eight gigabytes is workstation territory. Now you can run the biggest open-weight video models at full precision, serve 70B LLMs through vLLM, and start fine-tuning instead of just LoRA-ing. The RTX 5090, RTX A6000, and L40S are the typical homes for this tier.
e.g. NVIDIA RTX A6000 · NVIDIA L40 / L40S · Two 24 GB GPUs where the runtime supports sharding
Eighty gigabytes and up is datacenter territory — typically rented by the hour rather than owned. This is where you serve 70B+ LLMs at production scale, full fine-tune large models, and run multi-GPU video diffusion pipelines.
e.g. NVIDIA A100 80GB · NVIDIA H100 80GB · NVIDIA H200 141GB
When you don't have local hardware (or you need quality above what fits in your card), these cloud APIs cover the same workloads.
The benchmark commercial TTS / voice clone API.
Agentic coding in VS Code — reads, writes, runs, browses.
The benchmark for aesthetic image generation.
Run any open-source model with one API call.
Real-time inference platform — sub-second latency for diffusion.
Phone-scan to NeRF, Genie text-to-3D, and Dream Machine video.
The image model that actually renders text.
Cinematic-quality cloud video generation.
GPU marketplace — rent consumer cards at half the hyperscaler price.
Serverless GPU functions — deploy a Python file, get an HTTPS endpoint.
Cloud video gen with strong motion control.
Text-to-3D, image-to-3D, and texture generation for game pipelines.
Cloud video generation with Pikaffects and Scenes.