Anthropic
Claude Opus 4.8
Anthropic's flagship — top of the SWE-bench Verified leaderboard.
API onlytextimagefileREL 2026-05-27
Compare this model →Verified 2026-07-15
Overview
Claude Opus 4.8 is Anthropic's frontier flagship for deep reasoning and agentic coding. It currently leads SWE-bench Verified at 88.6% and tops the Artificial Analysis Intelligence Index, with a 1M-token context window and strong long-horizon agent behaviour. Premium pricing puts it in 'reach for it when it matters' territory.
Benchmark scores
Methodology →Intelligence61.4
Coding88.6
Reasoningn/a
Mathn/a
Sources and verification
- Coding: SWE-bench Verified · checked 2026-06-09
- Intelligence: Artificial Analysis Intelligence Index · checked 2026-07-15
- Catalogue and pricing: OpenRouter model record · checked 2026-07-15
Strengths and weaknesses
Strengths
- Best-in-class agentic coding
- 1M-token context
- Excellent instruction following on long tasks
Trade-offs
- Premium price ($5/$25 per 1M)
- Closed weights — API only
- Slower than mid-tier models
Reach for this model when
Hard coding tasks, large-repo refactors, and agent workflows where quality beats cost.
Run something like it locally
No open-weights equivalent. Closest local stand-ins are DeepSeek V4 or Kimi K2.6 served via vLLM — strong but a clear tier below on SWE-bench.