AI-Toolkit (Ostris)
Modern training framework — Flux, SDXL, SD3 LoRAs in YAML.
Sixteen gigabytes reduces compromise for multi-component image graphs and expands quantized LLM and video options. Unified memory on Apple Silicon is not directly interchangeable with discrete VRAM: the operating system and application share the same pool, and backend support differs.
These are planning envelopes. Model variant, precision, resolution, frame count, context, cache, runtime, and offload can move the same workload across tiers. The guide shows how to validate a real ComfyUI graph before buying hardware.
Sorted by best fit for this tier — tools designed around your VRAM budget first, then by our power-user score.
Modern training framework — Flux, SDXL, SD3 LoRAs in YAML.
Microsoft Research's structured 3D representation model.
Memory-efficient T2V via pyramidal flow matching.
The INRIA original — train your own splats.
Tencent's open 3D generator — multi-view, PBR, ready-to-use meshes.
Real-time-ish open video diffusion from Lightricks.
Stability's MMDiT flagship at 8B params.
Image-to-video diffusion — 25 frames, 14 or 25 steps.
Animation motion modules for ComfyUI.