Claude Fable 5 leads HLE (Humanity's Last Exam) in 2026 at 53.3%. Claude Opus 5 Max leads Intelligence Index at 60.7 (Xhigh 60.1, Fable 5 59.9). GPT-5.6 Sol max (HLE 47.2%, GPQA 94.1% co-leader), Claude Opus 4.8 (HLE 45.7%) compared. Ranked by GPQA Diamond, AIME, MATH-500, HLE. Free.
| # | Model | Score | Benchmarks | Input $/M | Output $/M | Speed | TTFT |
|---|---|---|---|---|---|---|---|
| 1 | 95 | 100 | $5.00 | $25.00 | 53 | 59.97s | |
| 2 | 94 | 99 | $5.00 | $25.00 | 54 | 40.50s | |
| 3 | 93 | 98 | $10.00 | $50.00 | 67 | 148.12s | |
| 4 | 93 | 97 | $5.00 | $25.00 | 52 | 11.81s | |
| 5 | Grok 4.6 (high) SpaceXAI | 92 | 96 | $2.00 | $6.00 | 56 | 36.89s |
| 6 | GPT-5.6 Sol (max) OpenAI | 92 | 96 | $5.00 | $30.00 | 65 | 208.71s |
| 7 | Kimi K3 (max) Kimi | 90 | 95 | $3.00 | $15.00 | 38 | 2.67s |
| 8 | GLM-5.3 (max) Z AI | 90 | 94 | $1.40 | $4.40 | 85 | 1.88s |
| 9 | GPT-5.6 Sol (xhigh) OpenAI | 89 | 93 | $5.00 | $30.00 | 63 | 50.77s |
| 10 | 88 | 93 | $5.00 | $25.00 | 53 | 4.92s | |
| 11 | Qwen3.8 Max Alibaba | 88 | 92 | $2.00 | $6.00 | 50 | 2.56s |
| 12 | Qwen3.8 2.4T A95B Alibaba | 87 | 91 | $2.00 | $6.00 | 45 | 2.53s |
| 13 | GPT-5.6 Sol (high) OpenAI | 87 | 91 | $5.00 | $30.00 | 70 | 9.27s |
| 14 | 87 | 91 | $5.00 | $25.00 | 56 | 12.05s | |
| 15 | GPT-5.6 Terra (max) OpenAI | 86 | 90 | $2.00 | $12.00 | 119 | 186.73s |
Models are scored using a weighted combination of benchmarks, pricing, and speed metrics relevant to this use case.
Models specifically designed for reasoning (like OpenAI o-series and DeepSeek R1) typically score highest on benchmarks like GPQA, AIME, and HLE. Check the rankings above for the latest results.
For tasks requiring genuine multi-step logic β math proofs, complex analysis, scientific research β yes. For simpler tasks, general-purpose models are more cost-effective.
Chain-of-thought (CoT) is when a model shows its step-by-step thinking process. Some models do this internally (hidden tokens), while others expose it. CoT generally improves accuracy on complex problems but increases token usage.