Fastest AI models ranked live by tokens/sec, TTFT latency, throughput. Celeris-1 leads raw speed at 2,121 tok/s (II 11.8); Mercury 2 fastest smart model at 891 tok/s (II 21.4); HyperNova 60B 2605 (417), Step 3.7 Flash (407, II 30.3), LFM2.5-VL-1.6B (386). Real-time rankings for 500+ models — Mercury 2, GPT-5.6, Claude Opus 5, Gemini 3, DeepSeek. Free.
| # | Model | Score | Benchmarks | Input $/M | Output $/M | Speed | TTFT |
|---|---|---|---|---|---|---|---|
| 1 | Celeris-1 Celeris | 84 | 18 | $0.20 | $0.70 | 1484 | 0.58s |
| 2 | Mercury 2 Inception | 67 | 36 | $0.25 | $0.75 | 835 | 3.93s |
| 3 | 61 | 88 | $0.75 | $3.75 | 317 | 4.63s | |
| 4 | Gemini 3.7 Flash (high) Google | 61 | 93 | $0.75 | $3.75 | 323 | 15.15s |
| 5 | Gemini 3.7 Flash (low) Google | 61 | 85 | $0.75 | $3.75 | 303 | 0.83s |
| 6 | 58 | 87 | $1.25 | $4.25 | 200 | 1.48s | |
| 7 | Qwen3.7 Max Alibaba | 56 | 78 | $2.50 | $7.50 | 212 | 2.24s |
| 8 | Gemini 3.6 Flash (high) Google | 56 | 85 | $0.75 | $3.75 | 222 | 15.93s |
| 9 | GLM-5.3 (max) Z AI | 56 | 95 | $1.40 | $4.40 | 85 | 1.88s |
| 10 | Gemini 3.5 Flash-Lite Google | 56 | 60 | $0.30 | $2.50 | 332 | 9.76s |
| 11 | 55 | 85 | $0.44 | $1.32 | 133 | 1.15s | |
| 12 | GLM-5.2 (max) Z AI | 55 | 85 | $1.40 | $4.40 | 113 | 1.72s |
| 13 | GPT-5.6 Sol (medium) OpenAI | 54 | 92 | $5.00 | $30.00 | 73 | 3.52s |
| 14 | Kimi K3 (max) Kimi | 54 | 96 | $3.00 | $15.00 | 38 | 2.67s |
| 15 | Qwen3.8 Max Alibaba | 54 | 92 | $2.00 | $6.00 | 50 | 2.56s |
Models are scored using a weighted combination of benchmarks, pricing, and speed metrics relevant to this use case.
Speed depends on the metric. Check the rankings above for both output speed (tokens/s) and TTFT (time to first token). Smaller models and those optimized for inference typically lead.
Not necessarily. Some smaller models achieve excellent benchmark scores while being much faster. The rankings above show both speed and quality so you can find the best tradeoff.
For a good user experience: TTFT under 1 second, output speed above 50 tok/s. For premium feel: TTFT under 0.3s, speed above 100 tok/s. Streaming helps mask latency.