Skip to content
xAI

Grok 4.20

2M-token context and aggressive pricing — xAI's speed-focused reasoner.

API onlytextimagefileREL 2026-03-31
Compare this model →Verified 2026-06-12

Overview

Grok 4.20 pairs the largest context window of any flagship-class model (2M tokens) with unusually low pricing at $1.25/$2.50 per 1M tokens. xAI emphasises agentic tool calling, speed, and a low hallucination rate. Independent benchmark coverage (benchlm.ai 72/100 provisional) places it solid-but-not-leading on raw intelligence.

Benchmark scores

Methodology →
Intelligence49.3
Codingn/a
Reasoningn/a
Mathn/a

Sources and verification

Strengths and weaknesses

Strengths

  • 2M-token context — largest in flagship class
  • Very cheap output tokens ($2.50/M)
  • Fast agentic tool calling

Trade-offs

  • Trails Opus/GPT-5.5 on hard reasoning
  • Closed weights

Reach for this model when

Huge-context analysis and high-volume agent loops on a budget.

Run something like it locally

Llama 4 Scout (10M context, open weights) is the long-context local alternative via vLLM.

vLLM

High-throughput LLM serving for GPUs.

OPEN SOURCE24–80 GB VRAM
VRAM fit24–80 GB