Welcome to the most comprehensive free AI benchmark leaderboard, tracking 500+ language models across 12 industry-standard benchmarks. Our data is updated hourly from Artificial Analysis, ensuring you always see the latest performance numbers for reasoning, coding, math, and general knowledge — including Anthropic reasoning eval benchmarks, the latest GPT-5.6 releases, and Gemini 3 Pro results.
Each model is evaluated on GPQA Diamond (graduate-level science reasoning), AIME 2025 (competition math), MMLU-Pro (multitask language understanding), HLE (hard logic and reasoning), LiveCodeBench and SWE-bench (coding ability), and the proprietary Intelligence Index. Unlike other leaderboards that only show scores, we pair every benchmark result with real-time per-token API pricing, so you can compare both performance and cost efficiency in one place.
Looking for the fastest AI model? Check the Speed Rankings for tokens-per-second and TTFT latency. Need reasoning power? The Math & Reasoning guide breaks down the top models by GPQA and AIME scores. Looking for LiveCodeBench 2026 top models? The interactive table below ranks all models by live LCB scores updated hourly. Comparing AI coding tools? See AI coding plans & subscriptions side-by-side. For the latest OpenAI API pricing August 2026 and all model costs, see our AI API Pricing Guide with live cost calculator. Need a quick estimate for your own token usage? Try the free AI API Cost Calculator — paste a prompt or enter tokens and see exact cost across GPT, Claude, Gemini, DeepSeek + 500 models.
The Anthropic reasoning eval is Anthropic's in-house evaluation suite for measuring how well Claude models reason through novel, open-ended problems that don't have a single correct answer. It complements industry benchmarks like GPQA Diamond, AIME 2025, and HLE by testing the kind of multi-step, real-world reasoning that production AI applications need — chain-of-thought robustness, refusal calibration, and the ability to recognize when a problem is underspecified.
Among the Claude models in the live Artificial Analysis Intelligence Index, Claude Opus 5 (Max) leads at II 60.7, followed by Claude Opus 5 (Xhigh) II 60.1 and Claude Fable 5 II 59.9. Our Claude family reasoning leaderboard pairs the AA Intelligence Index with GPQA Diamond, AIME 2025, HLE, and other public benchmarks across all Claude 4.x and 5.x models.
The leaderboard pairs Anthropic's internal reasoning eval with public benchmarks like GPQA Diamond, AIME 2025, MMLU-Pro, and HLE so visitors looking for one canonical reference can compare all major reasoning benchmarks on a single page. For the official Anthropic methodology and per-task breakdowns, see Anthropic's blog.
The eval typically covers multi-step reasoning, chain-of-thought robustness, and calibration — how consistently Claude models produce correct intermediate steps on novel, open-ended problems. Specific scoring methodologies and per-task breakdowns are published on Anthropic's blog. State-of-the-art Claude models (Opus 5, Fable 5) typically exceed public benchmark thresholds.
Live leaderboard across the benchmarks users ask about most — GPQA Diamond, AIME 2025, MMLU-Pro, HLE, LiveCodeBench and more. Claude Opus 5 (Adaptive Reasoning, Max Effort) leads the Intelligence Index at 60.7, followed by Opus 5 Xhigh (60.1) and Claude Fable 5 (59.9). Click any model for full pricing, speed, and side-by-side comparisons.
Compare 603+ AI models across 12 benchmarks — Intelligence, Coding, Math, Science, and more. Data updated hourly.
Composite score across math, science, coding
Graduate-level science Q&A (Diamond)
Knowledge & reasoning across 57 subjects
Live coding benchmark with new problems
American Invitational Math Exam
Competition-level math problems
Humanity's Last Exam - hardest questions
Composite coding benchmark score
Composite math benchmark score
Scientific coding problems
Instruction following benchmark
Terminal/CLI task completion
Compare pricing for all models side by side
Open AI API Cost Calculator →