OpenAI
GPT-5.5
OpenAI's flagship — state of the art on terminal/agentic benchmarks.
API onlytextimagefileREL 2026-04-24
Compare this model →Verified 2026-06-12
Overview
GPT-5.5 is OpenAI's frontier flagship and Opus 4.8's main rival. It leads Terminal-Bench 2.0 (82.7%) and FrontierMath, scores 82.6% on SWE-bench Verified (vals.ai) and 93.6% on GPQA Diamond, with a 1.05M-token context window at $5/$30 per 1M tokens.
Benchmark scores
Methodology →Intelligence60
Coding82.6
Reasoning93.6
Mathn/a
Sources and verification
- Coding: SWE-bench Verified · checked 2026-06-09
- Intelligence: Artificial Analysis Intelligence Index · checked 2026-06-12
- Reasoning: GPQA Diamond 93.6%
- Catalogue and pricing: OpenRouter model record · checked 2026-06-12
Strengths and weaknesses
Strengths
- SOTA on Terminal-Bench 2.0 (82.7%)
- Leads FrontierMath tiers 1–3
- Strong tool calling + agent autonomy
Trade-offs
- Output tokens pricier than Opus ($30 vs $25)
- Closed weights
Reach for this model when
Agentic coding, command-line automation, and frontier math/reasoning work.
Run something like it locally
No open equivalent. DeepSeek V4 Pro via vLLM is the strongest open-weights approximation.