Google
Gemini 3.1 Pro (Preview)
Google's reasoning flagship — top of the GPQA Diamond leaderboard.
API onlytextimagevideofileaudioREL 2026-02-19
Compare this model →Verified 2026-06-12
Overview
Gemini 3.1 Pro leads on science reasoning with 94.1–94.3% GPQA Diamond and sits at ~57 on the Artificial Analysis Intelligence Index. Full multimodal input and a 1M context window at $2/$12 per 1M tokens make it the cheapest of the true flagships.
Benchmark scores
Methodology →Intelligence57
Codingn/a
Reasoning95.5
Math98.1
Sources and verification
- Math: AIME · checked 2026-06-09
- Intelligence: Artificial Analysis Intelligence Index · checked 2026-06-12
- Reasoning: GPQA Diamond · checked 2026-06-09
- Catalogue and pricing: OpenRouter model record · checked 2026-06-12
Strengths and weaknesses
Strengths
- #1-tier GPQA Diamond science reasoning
- Cheapest true flagship ($2/$12)
- Video + audio input
Trade-offs
- Still 'preview' labelled
- 65K max output
Reach for this model when
Science/research reasoning, multimodal analysis, long-document work.
Run something like it locally
No open model matches its multimodal reasoning; DeepSeek V4 Pro is the text-only approximation.