February 2026 is preserved for reference. April 2026 is the latest published report.
Benchmark data verified . Page updated for series navigation.
February 2026
A comprehensive benchmark analysis for executives, founders and operators
What This Report Covers
Side-by-side benchmark results for 6 frontier AI models across reasoning, coding, agentic tasks, multimodal, long context, price, and speed. Sourced from official model cards and third-party evaluations as of February 27, 2026.
Bottom Line
No single model wins every category. Gemini 3.1 Pro leads reasoning. GPT-5.3 Codex leads autonomous coding. Claude Opus 4.6 leads agentic and computer-use tasks. Grok 4.1 leads speed and context. DeepSeek V3.2 leads cost and open-source access.
Quick Answers
Best for reasoning
Gemini 3.1 Pro
94.3% GPQA Diamond
Best for coding
GPT-5.3 Codex
77.3% Terminal-Bench 2.0
Best for agentic tasks
Claude Opus 4.6
53.0% HLE with Tools
Best value
DeepSeek V3.2
$0.28/M input tokens
Section 01
Gemini 3.1 Pro leads PhD-level and abstract reasoning in February 2026, scoring highest on GPQA Diamond (94.3%), ARC-AGI-2 (77.1%), and Humanity's Last Exam without tools (44.4%). Claude Opus 4.6 leads when tools are available, scoring 53.0% on HLE with Tools.
GPQA Diamond (PhD-level science)
ARC-AGI-2 (novel abstract reasoning)
Humanity's Last Exam (no tools)
HLE with Tools
AIME 2025 (competition math)
MMMLU (multilingual knowledge)
Section 02
GPT-5.3 Codex leads autonomous coding with 77.3% on Terminal-Bench 2.0, 56.8% on SWE-bench Pro, and a perfect 100% on AIME 2025. Claude Opus 4.6 leads computer-use with 72.7% on OSWorld-Verified. Gemini 3.1 Pro leads competitive programming at 75.6% on LiveCodeBench.
SWE-bench Verified (real GitHub bugs)
SWE-bench Pro (harder tasks)
Terminal-Bench 2.0 (agentic terminal)
OSWorld-Verified (computer-use agent)
LiveCodeBench (competitive programming)
Section 03
Claude Sonnet 4.6 leads office productivity with a GDPval-AA Elo of 1,633. Claude Opus 4.6 leads customer service agent tasks (93.5%) and finance agent benchmarks (60.7%). Gemini 3.1 Pro leads general web research at 85.9% on BrowseComp.
GDPval-AA Elo bars normalized to 1,700 max.
tau2-bench Retail (customer service agent)
GDPval-AA Elo (office productivity)
BrowseComp (web research agent)
Finance Agent
Section 04
Gemini 3.1 Pro leads vision and reasoning at 80.5% on MMMU-Pro. Claude Opus 4.6 and Gemini 3.1 Pro tie for 128K long-context retrieval at 84.9%. Claude Opus 4.6 leads extreme 1-million-token context at 76.0%. DeepSeek V3.2 is text-only with no vision support.
MMMU-Pro (vision + reasoning)
MRCR v2 128K (long-context retrieval)
MRCR v2 1M tokens (extreme long context)
Section 05
Full specifications for all six frontier models as of February 2026, including pricing, output speed, context window size, and the Artificial Analysis Intelligence Index composite score.
| Claude Opus 4.6 | Claude Sonnet 4.6 | Gemini 3.1 Pro | GPT-5.3 Codex | Grok 4.1 | DeepSeek V3.2 | |
|---|---|---|---|---|---|---|
| Released | Feb 5 '26 | Feb 17 '26 | Feb 19 '26 | Feb 5 '26 | Nov 17 '25 | Dec 1 '25 |
| Context Window | 200K (1Mβ) | 200K (1Mβ) | 1M std | ~200K | 2M | 128K |
| Input $/1M tokens | $5.00 | $3.00 | $2.00 | ~$1.75 | $0.20 | $0.28 |
| Output $/1M tokens | $25.00 | $15.00 | $12.00 | ~$14.00 | $0.50 | $0.42 |
| Speed (tokens/sec) | 67-72 | 54-56 | 91-110 | ~89 | 118 | 49 |
| Arena Elo | ~1,506 #1 | Top 5 | ~1,492 | Testing | ~1,483 | N/R |
| AA Intelligence Index | 53 | 52 | 57 | N/R | 35† | 32 |
| Open weights | No | No | No | No | No | MIT ✓ |
Section 06
Output cost and inference speed determine which model is practical at scale. Grok 4.1 leads speed at 118 tokens per second. DeepSeek V3.2 leads cost at $0.42 per million output tokens.
Output Price: $/1M Tokens
Speed: Tokens per Second
Section 07
Claude Opus 4.6
Best for complex agentic work
Claude Sonnet 4.6
Best production value
Gemini 3.1 Pro
Best reasoning + price-performance
GPT-5.3 Codex
Best autonomous coding
Grok 4.1
Best speed + value
DeepSeek V3.2
Best open-source
Original Analysis
Attainment's Price-Performance Index divides each model's AA Intelligence Index score by its output cost per million tokens. Higher is better. GPT-5.3 Codex is excluded: no AA Intelligence Index score is published as of this report.
| Model | AA Index | Output $/1M | Price-Performance |
|---|---|---|---|
| DeepSeek V3.2 | 32 | $0.42 | 76.2 pts/$ |
| Grok 4.1 | 35 | $0.50 | 70.0 pts/$ |
| Gemini 3.1 Pro | 57 | $12.00 | 4.8 pts/$ |
| Claude Sonnet 4.6 | 52 | $15.00 | 3.5 pts/$ |
| Claude Opus 4.6 | 53 | $25.00 | 2.1 pts/$ |
| GPT-5.3 Codex | N/A | $14.00 | N/A |
Price-Performance = AA Intelligence Index divided by output cost per million tokens. Higher score means more intelligence per dollar. Calculated by Attainment, February 2026.
Methodology
Benchmark scores were collected from official model cards, provider documentation, and independent evaluation platforms. All data reflects published results as of February 27, 2026. We selected benchmarks that measure distinct capabilities: PhD-level science knowledge (GPQA Diamond), abstract reasoning (ARC-AGI-2), real-world software engineering (SWE-bench), autonomous computing (OSWorld), and vision reasoning (MMMU-Pro). The AA Intelligence Index from Artificial Analysis provides a composite score normalized across evaluations.
Primary Sources
Summary
Gemini 3.1 Pro: Gemini 3.1 Pro leads raw reasoning: GPQA Diamond 94.3%, ARC-AGI-2 77.1%, AA Intelligence Index 57. Best reasoning-per-dollar of any frontier model in this report.
Claude Opus 4.6: Claude Opus 4.6 ranks highest by human preference (Arena Elo ~1,506) and leads agentic benchmarks: OSWorld 72.7%, HLE with Tools 53.0%, tau2-bench Retail 93.5%.
GPT-5.3 Codex: GPT-5.3 Codex leads autonomous coding: Terminal-Bench 2.0 at 77.3%, SWE-bench Pro at 56.8%, and a perfect 100% on AIME 2025 math.
Claude Sonnet 4.6: Claude Sonnet 4.6 is the best production value: leads office workflow Elo (1,633) and finance agent tasks (63.3%) at 60% of Claude Opus pricing.
Grok 4.1: Grok 4.1 is the fastest model (118 tokens per second) with the largest context window (2M tokens) and output pricing of $0.50/M: 50x cheaper than Claude Opus 4.6.
DeepSeek V3.2: DeepSeek V3.2 is the only MIT-licensed open-weight model, with competitive coding scores (LiveCodeBench 74.1%) and the cheapest API at $0.42/M output.
FAQ
Reading the model race because you have to deploy it?
Attainment helps owner-operated businesses turn frontier AI into production systems that lower costs and grow revenue. Strategy, AI automation, and the build, from one team.
Explore our AI automation servicesAttainment is a growth and operating-efficiency firm based in Canada. This report is produced independently with no paid placement or model provider sponsorship. All benchmark data is sourced from official model cards and third-party evaluations as of February 27, 2026.
Published ·Latest published edition
† Gemini MRCR 1M uses a different evaluation variant. † Grok 4.1 AA Index is for the Fast (Reasoning) variant.
N/R = not run. N/A = not applicable. ~ = estimated. ★ = category leader.
Sources: Anthropic, Google DeepMind, OpenAI, xAI, DeepSeek model cards; Artificial Analysis Intelligence Index v4.0; LMSYS Chatbot Arena; Vals AI; Digital Applied LLM Comparison. All scores as of Feb 27, 2026.