Fastest AI Models 2026 — Live Speed Leaderboard (Celeris-1, 2,121 tok/s)

Compare real-world speed for 311+ AI models: response latency, time to first token (TTFT), and throughput (tokens/sec). We pull live performance data from Artificial Analysis and pair it with per-token pricing so you can spot the fastest and most cost-efficient models for production APIs, real-time chat, and code completion.

For raw benchmark scores (GPQA, AIME, MMLU-Pro, HLE, SWE-bench) see the Benchmarks leaderboard. For pricing-only comparison try the LLM cost calculator.

Top 5 Fastest AI Models Right Now

  1. 1.Celeris-11483.7 tok/s
  2. 2.Mercury 2835.0 tok/s
  3. 3.Gemini 3.1 Flash-Lite346.2 tok/s
  4. 4.LFM2.5-8B-A1B338.6 tok/s
  5. 5.Gemini 3.5 Flash-Lite332.4 tok/s
Live speed data from Artificial Analysis API

AI Model Speed Rankings

Compare 311+ AI models by response speed, latency, and throughput. Find the fastest models for your use case.

311 models · click headers to sort
#
Model
Throughput
TTFT
$/1M
$/Speed
Price×TTFT
1
Celeris-1
Celeris
1484 t/s
580ms
$0.33
$0.000
188.5
2
Mercury 2
Inception
835 t/s
3.93s
$0.38
$0.000
1473.8
3
346 t/s
5.64s
$0.56
$0.002
3175.3
4
LFM2.5-8B-A1B
Liquid AI
339 t/s
1.95s
$0.00
5
332 t/s
9.76s
$0.85
$0.003
8296.0
6
326 t/s
1.04s
$0.13
$0.000
133.1
7
323 t/s
15.15s
$1.5
$0.005
22725.0
8
317 t/s
4.63s
$1.5
$0.005
6945.0
9
312 t/s
24.64s
$0.17
$0.001
4312.0
10
303 t/s
830ms
$1.5
$0.005
1245.0
11
302 t/s
1.16s
$0.11
$0.000
125.3
12
282 t/s
870ms
$0.07
$0.000
56.6
13
265 t/s
2.32s
$0.00
14
261 t/s
1.14s
$0.09
$0.000
100.3
15
255 t/s
1.84s
$0.28
$0.001
506.0
16
232 t/s
310ms
$0.17
$0.001
54.3
17
225 t/s
460ms
$0.85
$0.004
391.0
18
225 t/s
890ms
$0.09
$0.000
78.3
19
225 t/s
19.09s
$0.85
$0.004
16226.5
20
224 t/s
1.08s
$0.15
$0.001
162.0
21
222 t/s
15.93s
$1.5
$0.007
23895.0
22
o3-mini
OpenAI
221 t/s
6.45s
$1.9
$0.009
12416.3
23
220 t/s
20.91s
$1.9
$0.009
40251.8
24
214 t/s
1.14s
$0.15
$0.001
171.0
25
212 t/s
2.24s
$3.8
$0.018
8400.0
26
207 t/s
380ms
$0.00
27
205 t/s
610ms
$0.10
$0.000
61.0
28
200 t/s
1.48s
$2.0
$0.010
2960.0
29
198 t/s
2.25s
$0.41
$0.002
929.2
30
197 t/s
1.38s
$0.39
$0.002
542.3
31
195 t/s
730ms
$1.1
$0.006
821.3
32
189 t/s
6.57s
$1.1
$0.006
7391.3
33
186 t/s
4.08s
$0.07
$0.000
285.6
34
185 t/s
4.11s
$0.46
$0.003
1902.9
35
183 t/s
2.19s
$0.41
$0.002
904.5
36
180 t/s
20.69s
$3.4
$0.019
69828.8
37
176 t/s
2.85s
$0.46
$0.003
1319.5
38
Ling 3.0 Tiny
InclusionAI
175 t/s
2.57s
$0.00
39
171 t/s
37.51s
$0.14
$0.001
5176.4
40
171 t/s
27.29s
$3.4
$0.020
92103.7
41
170 t/s
1.08s
$0.10
$0.001
111.2
42
170 t/s
27.96s
$1.9
$0.011
53823.0
43
170 t/s
19.88s
$0.85
$0.005
16898.0
44
169 t/s
2.12s
$0.69
$0.004
1458.6
45
168 t/s
640ms
$0.46
$0.003
296.3
46
168 t/s
1.35s
$0.09
$0.001
118.8
47
167 t/s
94.31s
$0.14
$0.001
13014.8
48
163 t/s
11.89s
$1.7
$0.010
20070.3
49
163 t/s
1.73s
$0.30
$0.002
519.0
50
162 t/s
720ms
$1.7
$0.010
1215.4
51
161 t/s
860ms
$0.14
$0.001
118.7
52
161 t/s
880ms
$0.26
$0.002
230.6
53
160 t/s
890ms
$0.26
$0.002
231.4
54
160 t/s
800ms
$0.26
$0.002
209.6
55
160 t/s
8.12s
$1.7
$0.011
13706.6
56
159 t/s
159.40s
$0.45
$0.003
71730.0
57
158 t/s
7.39s
$0.85
$0.005
6281.5
58
156 t/s
10.98s
$0.85
$0.005
9333.0
59
154 t/s
810ms
$3.4
$0.022
2733.8
60
152 t/s
2.38s
$0.26
$0.002
623.6
61
Nova Lite
Amazon
151 t/s
960ms
$0.10
$0.001
100.8
62
151 t/s
2.06s
$0.69
$0.005
1417.3
63
151 t/s
860ms
$0.15
$0.001
129.0
64
150 t/s
780ms
$0.26
$0.002
204.4
65
149 t/s
810ms
$0.15
$0.001
121.5
66
146 t/s
43.84s
$0.45
$0.003
19728.0
67
146 t/s
14.42s
$0.85
$0.006
12257.0
68
146 t/s
1.66s
$1.1
$0.008
1889.1
69
146 t/s
20.94s
$1.6
$0.011
32729.2
70
146 t/s
1.08s
$0.85
$0.006
918.0
71
145 t/s
710ms
$0.45
$0.003
319.5
72
145 t/s
2.34s
$0.75
$0.005
1755.0
73
144 t/s
1.82s
$0.35
$0.002
637.0
74
144 t/s
11.08s
$0.45
$0.003
4986.0
75
144 t/s
2.14s
$0.45
$0.003
963.0
76
143 t/s
710ms
$0.17
$0.001
124.2
77
143 t/s
810ms
$0.30
$0.002
243.0
78
143 t/s
2.33s
$1.1
$0.008
2563.0
79
142 t/s
1.87s
$0.35
$0.002
654.5
80
142 t/s
800ms
$0.15
$0.001
120.0
81
GPT-4.1
OpenAI
142 t/s
870ms
$3.5
$0.025
3045.0
82
140 t/s
2.80s
$0.44
$0.003
1226.4
83
140 t/s
2.14s
$2.1
$0.015
4601.0
84
139 t/s
1.03s
$0.56
$0.004
579.9
85
Nex-N2-Pro
Nex AGI
138 t/s
1.61s
$1.0
$0.007
1610.0
86
138 t/s
930ms
$0.03
$0.000
26.0
87
138 t/s
1.74s
$0.45
$0.003
783.0
88
137 t/s
2.05s
$0.85
$0.006
1738.4
89
137 t/s
810ms
$0.26
$0.002
212.2
90
136 t/s
980ms
$0.09
$0.001
91.1
91
135 t/s
2.24s
$3.0
$0.022
6720.0
92
135 t/s
780ms
$0.30
$0.002
234.0
93
133 t/s
20.46s
$3.4
$0.026
70341.5
94
133 t/s
1.15s
$0.66
$0.005
759.0
95
132 t/s
64.93s
$4.8
$0.036
312508.1
96
132 t/s
2.28s
$3.0
$0.023
6840.0
97
132 t/s
2.25s
$0.00
98
131 t/s
1.01s
$4.4
$0.033
4418.8
99
131 t/s
144.08s
$5.6
$0.043
810450.0
100
130 t/s
2.33s
$1.1
$0.008
2563.0
Showing top 100 of 311 models. Use search/filter to narrow down.

Speed Metrics Guide

Throughput (tokens/s)

Output generation speed in tokens per second. Higher is better.

Good: >50 t/s · Excellent: >100 t/s
Time to First Token (TTFT)

Delay before the first token appears. Lower is better.

Good: <500ms · Excellent: <200ms
Price/Performance

Cost efficiency ratios. Lower values indicate better value.

$/Speed: price per t/s · Price×TTFT: latency penalty

Compare pricing for all models side by side

Open AI API Cost Calculator →