AI Benchmarks 2026 — Claude Opus 5 leads, Anthropic Reasoning Eval (500+ Models, Free & Live)

Welcome to the most comprehensive free AI benchmark leaderboard, tracking 500+ language models across 12 industry-standard benchmarks. Our data is updated hourly from Artificial Analysis, ensuring you always see the latest performance numbers for reasoning, coding, math, and general knowledge — including Anthropic reasoning eval benchmarks, the latest GPT-5.6 releases, and Gemini 3 Pro results.

Each model is evaluated on GPQA Diamond (graduate-level science reasoning), AIME 2025 (competition math), MMLU-Pro (multitask language understanding), HLE (hard logic and reasoning), LiveCodeBench and SWE-bench (coding ability), and the proprietary Intelligence Index. Unlike other leaderboards that only show scores, we pair every benchmark result with real-time per-token API pricing, so you can compare both performance and cost efficiency in one place.

Looking for the fastest AI model? Check the Speed Rankings for tokens-per-second and TTFT latency. Need reasoning power? The Math & Reasoning guide breaks down the top models by GPQA and AIME scores. Looking for LiveCodeBench 2026 top models? The interactive table below ranks all models by live LCB scores updated hourly. Comparing AI coding tools? See AI coding plans & subscriptions side-by-side. For the latest OpenAI API pricing August 2026 and all model costs, see our AI API Pricing Guide with live cost calculator. Need a quick estimate for your own token usage? Try the free AI API Cost Calculator — paste a prompt or enter tokens and see exact cost across GPT, Claude, Gemini, DeepSeek + 500 models.

Benchmarks Tracked (August 2026)

What is the Anthropic Reasoning Eval? (Claude family reasoning benchmark)

The Anthropic reasoning eval is Anthropic's in-house evaluation suite for measuring how well Claude models reason through novel, open-ended problems that don't have a single correct answer. It complements industry benchmarks like GPQA Diamond, AIME 2025, and HLE by testing the kind of multi-step, real-world reasoning that production AI applications need — chain-of-thought robustness, refusal calibration, and the ability to recognize when a problem is underspecified.

Among the Claude models in the live Artificial Analysis Intelligence Index, Claude Opus 5 (Max) leads at II 60.7, followed by Claude Opus 5 (Xhigh) II 60.1 and Claude Fable 5 II 59.9. Our Claude family reasoning leaderboard pairs the AA Intelligence Index with GPQA Diamond, AIME 2025, HLE, and other public benchmarks across all Claude 4.x and 5.x models.

The leaderboard pairs Anthropic's internal reasoning eval with public benchmarks like GPQA Diamond, AIME 2025, MMLU-Pro, and HLE so visitors looking for one canonical reference can compare all major reasoning benchmarks on a single page. For the official Anthropic methodology and per-task breakdowns, see Anthropic's blog.

How is the Anthropic Reasoning Eval scored?

The eval typically covers multi-step reasoning, chain-of-thought robustness, and calibration — how consistently Claude models produce correct intermediate steps on novel, open-ended problems. Specific scoring methodologies and per-task breakdowns are published on Anthropic's blog. State-of-the-art Claude models (Opus 5, Fable 5) typically exceed public benchmark thresholds.

Top Performing AI Models (August 2026)

Live leaderboard across the benchmarks users ask about most — GPQA Diamond, AIME 2025, MMLU-Pro, HLE, LiveCodeBench and more. Claude Opus 5 (Adaptive Reasoning, Max Effort) leads the Intelligence Index at 60.7, followed by Opus 5 Xhigh (60.1) and Claude Fable 5 (59.9). Click any model for full pricing, speed, and side-by-side comparisons.

Live data from Artificial Analysis API

AI Model Benchmarks

Compare 603+ AI models across 12 benchmarks — Intelligence, Coding, Math, Science, and more. Data updated hourly.

Benchmarks:
603 models · click headers to sort
#
Model
Speed
$/1M
AA Index
GPQA
MMLU-Pro
LiveCode
AIME
HLE
1
53 t/s
$10.0
63.1
2
54 t/s
$10.0
62.5
4
52 t/s
$10.0
61.5
5
56 t/s
$3.0
60.9
6
65 t/s
$11.3
60.9
7
38 t/s
$6.0
59.7
8
85 t/s
$2.1
59.5
9
63 t/s
$11.3
59.0
10
53 t/s
$10.0
58.6
11
50 t/s
$3.0
58.1
12
45 t/s
$3.0
57.7
13
70 t/s
$11.3
57.3
14
56 t/s
$10.0
57.3
15
$2.0
56.8
16
119 t/s
$4.5
56.6
17
85 t/s
$11.3
56.3
18
323 t/s
$1.5
56.0
19
64 t/s
$3.0
55.8
20
73 t/s
$11.3
55.6
21
76 t/s
$4.0
55.3
22
48 t/s
$10.0
55.0
23
74 t/s
$11.3
54.7
24
317 t/s
$1.5
53.4
25
74 t/s
$2.0
53.2
26
200 t/s
$2.0
53.2
27
131 t/s
$5.6
53.1
28
101 t/s
$4.5
52.8
29
113 t/s
$2.1
52.6
30
49 t/s
$10.0
52.5
31
159 t/s
$0.45
52.3
32
171 t/s
$3.4
52.0
33
$1.1
52.0
34
133 t/s
$0.66
51.8
35
222 t/s
$1.5
51.6
36
76 t/s
$11.3
51.4
37
303 t/s
$1.5
50.9
38
57 t/s
$11.3
50.7
39
103 t/s
$4.5
50.1
40
146 t/s
$0.45
50.1
41
55 t/s
$6.0
48.4
42
36 t/s
$6.0
48.3
43
108 t/s
$4.5
47.7
44
Motif 3
Motif Technologies
47.4
45
144 t/s
$0.45
47.0
46
98 t/s
$4.5
46.8
47
180 t/s
$3.4
46.7
48
212 t/s
$3.8
46.7
49
132 t/s
$4.8
45.5
50
MiniMax-M3
MiniMax
90 t/s
$0.53
45.4
51
74 t/s
$0.54
45.3
52
46 t/s
$1.7
45.1
53
40 t/s
$10.0
44.9
54
65 t/s
$1.1
44.5
55
74 t/s
$11.3
44.5
56
44.3
57
44 t/s
$10.0
43.9
58
76 t/s
$0.54
43.7
59
78 t/s
$4.8
43.3
60
49 t/s
$1.7
43.0
61
59 t/s
$0.54
42.9
62
65 t/s
$1.1
42.9
63
58 t/s
$4.0
42.6
64
Inkling (xhigh)
Thinking Machines
74 t/s
$1.8
42.3
65
Hy3
Tencent
67 t/s
$0.24
42.2
66
$0.17
42.1
67
46 t/s
$10.0
41.9
68
67 t/s
$11.3
41.9
69
Nex-N2-Pro
Nex AGI
138 t/s
$1.0
41.7
70
47 t/s
$0.53
41.6
71
41.4
72
94 t/s
$4.5
41.3
73
$4.8
41.2
74
Inkling Small
Thinking Machines
127 t/s
$0.53
41.2
75
61 t/s
$2.9
41.1
76
79 t/s
$2.1
41.0
77
163 t/s
$1.7
40.9
78
55 t/s
$1.3
40.7
79
46 t/s
$1.6
40.6
80
$4.5
40.6
81
56 t/s
$1.1
40.5
82
99 t/s
$5.6
40.2
83
39.9
84
130 t/s
$0.56
39.7
85
185 t/s
$0.46
39.7
86
57 t/s
$0.70
39.4
87
39.1
88
$0.17
39.0
89
61 t/s
$0.53
38.9
90
144 t/s
$0.45
38.9
91
$4.8
38.9
92
36 t/s
$10.0
38.8
93
189 t/s
$1.1
38.7
94
146 t/s
$1.1
38.3
95
MiMo-V2.5
Xiaomi
56 t/s
$0.17
38.0
96
99 t/s
$1.6
38.0
97
146 t/s
$1.6
37.9
98
Ling 3.0 Flash
InclusionAI
$0.11
37.8
99
58 t/s
$1.4
37.7
100
112 t/s
$3.4
37.5
Showing top 100 of 603 models. Use search/filter to narrow down.

Benchmark Guide

Intelligence
Source ↗

Composite score across math, science, coding

Graduate-level science Q&A (Diamond)

MMLU-Pro
Source ↗

Knowledge & reasoning across 57 subjects

LiveCodeBench
Source ↗

Live coding benchmark with new problems

AIME 2025
Source ↗

American Invitational Math Exam

MATH-500
Source ↗

Competition-level math problems

Humanity's Last Exam - hardest questions

Composite coding benchmark score

Composite math benchmark score

SciCode
Source ↗

Scientific coding problems

IFBench
Source ↗

Instruction following benchmark

TerminalBench
Source ↗

Terminal/CLI task completion

Compare pricing for all models side by side

Open AI API Cost Calculator →