Advertisement

Fastest AI Models 2026 — Live Speed Leaderboard (Celeris-1, 2,121 tok/s)

Compare real-world speed for 329+ AI models: response latency, time to first token (TTFT), and throughput (tokens/sec). We pull live performance data from Artificial Analysis and pair it with per-token pricing so you can spot the fastest and most cost-efficient models for production APIs, real-time chat, and code completion.

For raw benchmark scores (GPQA, AIME, MMLU-Pro, HLE, SWE-bench) see the Benchmarks leaderboard. For pricing-only comparison try the LLM cost calculator.

Top 5 Fastest AI Models Right Now

  1. 1.Celeris-11459.7 tok/s
  2. 2.Mercury 2880.8 tok/s
  3. 3.Gemini 2.5 Flash-Lite (Reasoning)415.9 tok/s
  4. 4.Gemini 3.5 Flash-Lite365.0 tok/s
  5. 5.HyperNova 60B 2605 (high, based on gpt-oss-120b)353.5 tok/s
Live speed data from Artificial Analysis API

AI Model Speed Rankings

Compare 329+ AI models by response speed, latency, and throughput. Find the fastest models for your use case.

329 models · click headers to sort
#
Model
Throughput ↓
TTFT
$/1M
$/Speed
Price×TTFT
1
Celeris-1
Celeris
1460 t/s
600ms
$0.33
$0.000
195.0
2
Mercury 2
Inception
881 t/s
4.60s
$0.38
$0.000
1725.0
3
416 t/s
25.55s
$0.17
$0.000
4471.2
4
365 t/s
8.61s
$0.85
$0.002
7318.5
5
353 t/s
600ms
$0.07
$0.000
39.0
6
Ling 3.0 Flash
InclusionAI
323 t/s
2.78s
$0.11
$0.000
300.2
7
319 t/s
1.31s
$0.41
$0.001
541.0
8
308 t/s
810ms
$1.5
$0.005
1215.0
9
303 t/s
16.28s
$1.5
$0.005
24420.0
10
302 t/s
570ms
$0.10
$0.000
54.1
11
298 t/s
5.04s
$0.56
$0.002
2837.5
12
298 t/s
280ms
$0.17
$0.001
49.0
13
289 t/s
4.96s
$1.5
$0.005
7440.0
14
288 t/s
10.22s
$1.5
$0.005
15330.0
15
287 t/s
890ms
$0.07
$0.000
57.9
16
281 t/s
850ms
$0.10
$0.000
87.5
17
246 t/s
390ms
$0.00
—
—
18
240 t/s
810ms
$0.25
$0.001
200.9
19
Agnes 3.0 Flash
Sapiens AI
237 t/s
1.84s
$0.07
$0.000
138.0
20
234 t/s
131.61s
$1.7
$0.007
222157.7
21
233 t/s
23.48s
$2.0
$0.009
46960.0
22
231 t/s
760ms
$0.09
$0.000
70.7
23
230 t/s
840ms
$0.26
$0.001
218.4
24
227 t/s
1.90s
$0.28
$0.001
522.5
25
226 t/s
13.39s
$3.4
$0.015
45191.3
26
225 t/s
21.00s
$1.9
$0.009
40425.0
27
223 t/s
1.16s
$0.09
$0.000
102.1
28
223 t/s
1.19s
$0.53
$0.002
624.8
29
220 t/s
22.48s
$0.85
$0.004
19108.0
30
217 t/s
22.70s
$2.0
$0.009
45400.0
31
216 t/s
18.36s
$3.4
$0.016
61965.0
32
214 t/s
1.32s
$0.66
$0.003
871.2
33
214 t/s
1.43s
$0.66
$0.003
943.8
34
214 t/s
5.93s
$1.1
$0.005
6671.3
35
214 t/s
610ms
$0.10
$0.000
61.0
36
214 t/s
420ms
$0.05
$0.000
22.3
37
207 t/s
840ms
$3.4
$0.016
2835.0
38
205 t/s
940ms
$1.1
$0.005
1057.5
39
204 t/s
1.99s
$1.1
$0.005
2089.5
40
201 t/s
2.68s
$3.8
$0.019
10050.0
41
198 t/s
1.16s
$0.30
$0.002
348.0
42
197 t/s
2.38s
$0.41
$0.002
982.9
43
o3-mini
OpenAI
196 t/s
5.09s
$1.9
$0.010
9798.3
44
194 t/s
430ms
$0.85
$0.004
365.5
45
193 t/s
19.50s
$1.5
$0.008
29250.0
46
187 t/s
720ms
$0.17
$0.001
126.0
47
184 t/s
6.15s
$1.7
$0.009
10381.2
48
181 t/s
18.73s
$2.0
$0.011
37460.0
49
181 t/s
2.24s
$0.41
$0.002
925.1
50
Inkling Small
Thinking Machines
178 t/s
2.13s
$0.53
$0.003
1118.3
51
178 t/s
4.07s
$0.56
$0.003
2291.4
52
177 t/s
970ms
$0.09
$0.000
85.4
53
177 t/s
81.70s
$0.14
$0.001
11274.6
54
Nova Lite
Amazon
176 t/s
910ms
$0.10
$0.001
95.5
55
176 t/s
14.98s
$0.85
$0.005
12733.0
56
174 t/s
830ms
$0.14
$0.001
114.5
57
174 t/s
1.02s
$0.90
$0.005
918.0
58
173 t/s
2.39s
$0.30
$0.002
728.9
59
171 t/s
4.76s
$0.46
$0.003
2203.9
60
171 t/s
840ms
$0.46
$0.003
388.9
61
Ling 3.0 Tiny
InclusionAI
171 t/s
2.71s
$0.00
—
—
62
170 t/s
95.64s
$0.46
$0.003
44281.3
63
169 t/s
830ms
$1.7
$0.010
1401.0
64
169 t/s
760ms
$0.26
$0.002
199.1
65
168 t/s
8.33s
$0.85
$0.005
7080.5
66
166 t/s
10.55s
$0.85
$0.005
8967.5
67
166 t/s
34.91s
$0.14
$0.001
4817.6
68
165 t/s
2.08s
$0.69
$0.004
1431.0
69
GPT-4.1
OpenAI
165 t/s
960ms
$3.5
$0.021
3360.0
70
161 t/s
720ms
$0.26
$0.002
188.6
71
160 t/s
1.77s
$0.09
$0.001
155.8
72
159 t/s
780ms
$0.30
$0.002
234.0
73
157 t/s
2.40s
$0.75
$0.005
1800.0
74
157 t/s
940ms
$4.4
$0.028
4112.5
75
156 t/s
1.58s
$2.1
$0.014
3397.0
76
155 t/s
750ms
$0.26
$0.002
196.5
77
155 t/s
740ms
$0.15
$0.001
111.0
78
155 t/s
1.07s
$0.85
$0.006
909.5
79
155 t/s
2.48s
$0.26
$0.002
649.8
80
154 t/s
870ms
$0.26
$0.002
227.9
81
154 t/s
750ms
$0.15
$0.001
112.5
82
154 t/s
800ms
$0.15
$0.001
120.0
83
151 t/s
27.93s
$1.9
$0.013
53765.3
84
148 t/s
3.43s
$0.15
$0.001
514.5
85
148 t/s
2.04s
$0.69
$0.005
1403.5
86
147 t/s
2.27s
$3.0
$0.020
6810.0
87
146 t/s
3.55s
$0.15
$0.001
532.5
88
145 t/s
93.69s
$5.6
$0.039
527006.3
89
145 t/s
2.37s
$0.00
—
—
90
o3
OpenAI
144 t/s
4.51s
$3.5
$0.024
15785.0
91
144 t/s
3.02s
$0.40
$0.003
1208.0
92
143 t/s
1.84s
$0.35
$0.002
644.0
93
142 t/s
2.23s
$3.0
$0.021
6690.0
94
142 t/s
2.20s
$0.80
$0.006
1760.0
95
141 t/s
2.23s
$0.80
$0.006
1784.0
96
Devstral 2
Mistral
141 t/s
2.25s
$0.00
—
—
97
139 t/s
840ms
$0.03
$0.000
23.5
98
139 t/s
2.23s
$0.00
—
—
99
137 t/s
3.82s
$0.30
$0.002
1146.0
100
137 t/s
2.05s
$0.85
$0.006
1738.4
Showing top 100 of 329 models. Use search/filter to narrow down.

Speed Metrics Guide

Throughput (tokens/s)

Output generation speed in tokens per second. Higher is better.

Good: >50 t/s · Excellent: >100 t/s
Time to First Token (TTFT)

Delay before the first token appears. Lower is better.

Good: <500ms · Excellent: <200ms
Price/Performance

Cost efficiency ratios. Lower values indicate better value.

$/Speed: price per t/s · Price×TTFT: latency penalty

Compare pricing for all models side by side

Open AI API Cost Calculator →