March 2026 is preserved for reference. April 2026 is the latest published report.
Benchmark data verified . Page updated for series navigation.
March 2026
8 frontier models compared across reasoning, coding, agentic tasks, multimodal, speed, and pricing. For executives, founders, and operators.
What This Report Covers
Side-by-side benchmark results for 8 frontier AI models across reasoning, coding, agentic tasks, multimodal, long context, price, and speed. This edition adds GPT-5.4, Grok 4.20 Beta, MiniMax M2.7, and GLM-5 to the comparison. Sourced from official model cards and third-party evaluations as of March 20, 2026.
Bottom Line
No single model wins every category. GPT-5.4 is the biggest mover, surpassing human experts on OSWorld. Gemini 3.1 Pro holds the reasoning crown. Claude Opus 4.6 leads human preference and agentic work. MiniMax M2.7 and GLM-5 prove frontier intelligence is no longer exclusive to Western labs or closed models.
What Changed This Month
View February edition →GPT-5.4 surpasses human experts
First model to beat human baseline on OSWorld (75.0% vs 72.4%). 1.05M context window. Replaces GPT-5.3 Codex in this report.
Grok 4.20 Beta: 212+ tokens/sec
Nearly 2x the speed of any competitor. Non-hallucination record of 78%. 2M token context. Replaces Grok 4.1.
MiniMax M2.7 enters the report
AA Intelligence Index 50 at $1.20/M output. Self-evolving model that handled 30-50% of its own RL research.
GLM-5 enters as open-source leader
First open-weight model to hit AA Index 50. 744B/40B MoE, MIT license. Trained entirely on Huawei chips.
Quick Answers
Best for reasoning
Gemini 3.1 Pro
94.3% GPQA Diamond
Best for coding
GPT-5.4
75.0% OSWorld (beats humans)
Best for agentic tasks
Claude Opus 4.6
53.0% HLE with Tools
Best value
MiniMax M2.7
AA Index 50, $1.20/M output
Section 01
Gemini 3.1 Pro leads PhD-level science and abstract reasoning in March 2026, scoring highest on GPQA Diamond (94.3%) and ARC-AGI-2 (77.1%). GPT-5.4 is a strong second at 92.8% GPQA Diamond and 73.3% ARC-AGI-2. On Humanity's Last Exam with tools, Claude Opus 4.6 leads at 53.0%.
GPQA Diamond (PhD-level science)
ARC-AGI-2 (novel abstract reasoning)
Humanity's Last Exam (no tools)
HLE with Tools
AIME 2025 (competition math)
MMMLU (multilingual knowledge)
Section 02
GPT-5.4 is the new coding leader in March 2026, surpassing human experts on OSWorld (75.0% vs 72.4% human baseline) and leading SWE-bench Pro at 57.7%. Claude Opus 4.6 holds the top SWE-bench Verified score at 80.8%. GLM-5, the open-source leader, scores 77.8% on SWE-bench Verified.
SWE-bench Verified (real GitHub bugs)
SWE-bench Pro (harder tasks)
Terminal-Bench 2.0 (agentic terminal)
OSWorld-Verified (computer-use agent)
LiveCodeBench (competitive programming)
Section 03
GPT-5.4 leads office productivity with a GDPval-AA score of 1,667 and web research at 82.7% on BrowseComp. Claude Opus 4.6 leads customer service agent tasks (93.5% tau2-bench Retail). GLM-5 brings open-source agentic capabilities with 89.7% on tau2-bench.
GDPval-AA bars normalized to 1,700 max.
tau2-bench Retail (customer service agent)
GDPval-AA (office productivity)
BrowseComp (web research agent)
Finance Agent
Section 04
GPT-5.4 leads vision reasoning at 90.6% on MMMU-Pro (Pro variant). Claude Opus 4.6 leads extreme 1-million-token context at 76.0% on MRCR v2. DeepSeek V3.2 and GLM-5 are text-only with no vision support. MiniMax M2.7 scores 83.5% on MMMU.
MMMU-Pro (vision + reasoning)
MRCR v2 128K (long-context retrieval)
MRCR v2 1M tokens (extreme long context)
Section 05
Full specifications for all eight frontier models as of March 20, 2026, including pricing, output speed, context window size, and the Artificial Analysis Intelligence Index composite score.
| Opus 4.6 | Sonnet 4.6 | Gemini 3.1 | GPT-5.4 | Grok 4.20 | DeepSeek | MiniMax | GLM-5 | |
|---|---|---|---|---|---|---|---|---|
| Released | Feb 5 '26 | Feb 17 '26 | Feb 19 '26 | Mar 5 '26 | Mar 10 '26 | Dec 1 '25 | Mar 18 '26 | Feb 11 '26 |
| Context Window | 200K (1Mβ) | 200K (1Mβ) | 1M std | 1.05M | 2M | 128K | ~200K | 200K |
| Input $/1M tokens | $5.00 | $3.00 | $2.00 | $2.50 | $2.00 | $0.28 | $0.30 | ~$1.00 |
| Output $/1M tokens | $25.00 | $15.00 | $12.00 | $15.00 | $6.00 | $0.42 | $1.20 | ~$3.20 |
| Speed (tokens/sec) | 67-72 | 54-56 | ~120 | ~82 | 212+ | 49 | ~49 | ~85 |
| Arena Elo | ~1,501 | Top 5 | ~1,493 | ~1,480† | ~1,493 | N/R | 1,407† | 1,455 |
| AA Intelligence Index | 53 | 52 | 57 | 57 | 48 | 32 | 50 | 50 |
| Open weights | No | No | No | No | No | MIT ✓ | No | MIT ✓ |
Section 06
Output cost and inference speed determine which model is practical at scale. Grok 4.20 Beta leads speed at 212+ tokens per second. DeepSeek V3.2 leads cost at $0.42 per million output tokens. MiniMax M2.7 offers the best balance of intelligence and price.
Output Price: $/1M Tokens
Speed: Tokens per Second
Section 07
Claude Opus 4.6
Best for complex agentic work
Claude Sonnet 4.6
Best production value (Anthropic)
Gemini 3.1 Pro
Best reasoning + price-performance
GPT-5.4
Best for autonomous tasks
Grok 4.20 Beta
Fastest frontier model
DeepSeek V3.2
Cheapest API
MiniMax M2.7
Best value frontier intelligence
GLM-5
Best open-source model
Original Analysis
Attainment's Price-Performance Index divides each model's AA Intelligence Index score by its output cost per million tokens. Higher is better. MiniMax M2.7 enters the report at 41.7 pts/$, the second-best ratio behind DeepSeek V3.2.
| Model | AA Index | Output $/1M | Price-Performance |
|---|---|---|---|
| DeepSeek V3.2 | 32 | $0.42 | 76.2 pts/$ |
| MiniMax M2.7 | 50 | $1.20 | 41.7 pts/$ |
| GLM-5 | 50 | $3.20 | 15.6 pts/$ |
| Grok 4.20 Beta | 48 | $6.00 | 8.0 pts/$ |
| Gemini 3.1 Pro | 57 | $12.00 | 4.8 pts/$ |
| GPT-5.4 | 57 | $15.00 | 3.8 pts/$ |
| Claude Sonnet 4.6 | 52 | $15.00 | 3.5 pts/$ |
| Claude Opus 4.6 | 53 | $25.00 | 2.1 pts/$ |
Price-Performance = AA Intelligence Index divided by output cost per million tokens. Higher score means more intelligence per dollar. Calculated by Attainment, March 2026.
Methodology
Benchmark scores were collected from official model cards, provider documentation, and independent evaluation platforms. All data reflects published results as of March 20, 2026. We selected benchmarks that measure distinct capabilities: PhD-level science knowledge (GPQA Diamond), abstract reasoning (ARC-AGI-2), real-world software engineering (SWE-bench), autonomous computing (OSWorld), and vision reasoning (MMMU-Pro). The AA Intelligence Index from Artificial Analysis provides a composite score normalized across evaluations. This is the second edition in a monthly series.
Primary Sources
Summary
GPT-5.4: GPT-5.4 is the biggest mover this month. It surpasses human experts on OSWorld (75.0% vs 72.4%), leads SWE-bench Pro (57.7%), Terminal-Bench 2.0 (75.1%), and GDPval-AA (1,667). AA Intelligence Index ties Gemini at 57.
Gemini 3.1 Pro: Gemini 3.1 Pro holds the reasoning crown: GPQA Diamond 94.3%, ARC-AGI-2 77.1%, AA Intelligence Index 57. At $12/M output, it remains the best reasoning-per-dollar among top-tier models.
Claude Opus 4.6: Claude Opus 4.6 retains the human preference lead (Arena Elo ~1,501) and leads agentic benchmarks: HLE with Tools 53.0%, tau2-bench Retail 93.5%, SWE-bench Verified 80.8%.
MiniMax M2.7: MiniMax M2.7 is the value story of the month. AA Intelligence Index 50 at $1.20/M output gives it 41.7 pts/$ price-performance. It autonomously handled 30-50% of its own reinforcement learning research.
GLM-5: GLM-5 is the first open-weight model to hit 50 on the AA Intelligence Index. SWE-bench Verified 77.8%, HLE with Tools 50.4%, MIT license. Trained entirely on Huawei Ascend chips.
Grok 4.20 Beta: Grok 4.20 Beta is the fastest at 212+ tokens/sec with the largest context (2M tokens). Its 78% non-hallucination rate sets a new record. At $6/M output, it trades raw intelligence for speed and honesty.
Claude Sonnet 4.6: Claude Sonnet 4.6 is the quiet workhorse: Finance Agent leader (63.3%), OSWorld 72.5%, SWE-bench Verified 79.6%, all at 60% of Opus pricing.
DeepSeek V3.2: DeepSeek V3.2 remains the cheapest API at $0.42/M output, with strong math (AIME 96.0%) and coding (LiveCodeBench 74.1%). V4 is expected in April.
FAQ
Reading the model race because you have to deploy it?
Attainment helps owner-operated businesses turn frontier AI into production systems that lower costs and grow revenue. Strategy, AI automation, and the build, from one team.
Explore our AI automation services† GPT-5.4 MMMU-Pro score is from the Pro variant ($30/M input). GPT-5.4 and MiniMax M2.7 Arena Elo scores are preliminary.
N/R = not run. N/A = not applicable. ~ = estimated. ★ = category leader.
Sources: Anthropic, Google DeepMind, OpenAI, xAI, DeepSeek, MiniMax, Zhipu AI model cards; Artificial Analysis Intelligence Index v4.0; Arena.ai (formerly LMSYS); Vals AI; Digital Applied LLM Comparison. All scores as of March 20, 2026.