PinchBench evaluates AI models on 6 real-world agent scenarios: coding agent (LiveCodeBench, TerminalBench, SciCode), reasoning & logic (GPQA, AIME 2025, MATH-500, HLE), instruction following (IFBench, MMLU-Pro), research & analysis, and tool use & agentic (τ²-bench, LCR). For raw academic scores, see our full LLM benchmark leaderboard, and for raw speed numbers, the AI model speed rankings.
PinchBench is a real-world AI agent benchmark that scores 400+ models on production-style agent tasks — coding agent scenarios (LiveCodeBench, TerminalBench, SciCode), reasoning & logic (GPQA Diamond, AIME 2025, HLE, MATH-500), instruction following (IFBench, MMLU-Pro), tool use & agentic (τ²-bench, LCR), and research workflows. Each model receives two composite scores: a Coding Index weighted toward agentic coding performance and an Intelligence Index weighted toward general reasoning.
Live data · Last updated · Refreshed hourly from Artificial Analysis
Top agents today by Coding Index: GPT-5.6 Sol (xhigh) 78.3, Claude Opus 5 (Max) 78.0, GPT-5.6 Sol (max) 77.4. Top agents by Intelligence Index: Claude Opus 5 (Max) II 60.7, Claude Opus 5 (Xhigh) II 60.1, Claude Fable 5 II 59.9. The leaderboard is updated hourly from Artificial Analysis data — no signup required.
Most benchmarks measure single-turn academic performance (GPQA, MMLU-Pro, AIME 2025). PinchBench measures end-to-end agent behavior: can the model write a function, run it, see it fail, and iterate? That requires tool use (τ²-bench, LCR), instruction following (IFBench), and multi-step coding (LiveCodeBench, TerminalBench, SciCode). It's the difference between "can solve a math problem" and "can build a working app from a spec."
How do AI models perform on real agent tasks? PinchBench scores 603+ models across coding, reasoning, tool use, and instruction following — with live pricing data.
Balanced score across all agent capabilities
| # | Model | Score | Bar | Input $/M | Output $/M | Speed | TTFT | Efficiency |
|---|---|---|---|---|---|---|---|---|
| 1 | Claude Opus 5 (Adaptive Reasoning, Max Effort) Anthropic | 70.5 | $5.00 | $25.00 | 53 | 59.97s | 7.1 | |
| 2 | Claude Opus 5 (Adaptive Reasoning, Xhigh Effort) Anthropic | 69.8 | $5.00 | $25.00 | 54 | 40.50s | 7.0 | |
| 3 | Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback) Anthropic | 69.3 | $10.00 | $50.00 | 67 | 148.12s | 3.5 | |
| 4 | GPT-5.6 Sol (max) OpenAI | 69.2 | $5.00 | $30.00 | 65 | 208.71s | 6.1 | |
| 5 | Claude Opus 5 (Adaptive Reasoning, High Effort) Anthropic | 69.0 | $5.00 | $25.00 | 52 | 11.81s | 6.9 | |
| 6 | Grok 4.6 (high) SpaceXAI | 68.9 | $2.00 | $6.00 | 56 | 36.89s | 23.0 | |
| 7 | GPT-5.6 Sol (xhigh) OpenAI | 68.6 | $5.00 | $30.00 | 63 | 50.77s | 6.1 | |
| 8 | Kimi K3 (max) Kimi | 68.0 | $3.00 | $15.00 | 38 | 2.67s | 11.3 | |
| 9 | GPT-5.6 Sol (high) OpenAI | 67.3 | $5.00 | $30.00 | 70 | 9.27s | 6.0 | |
| 10 | GLM-5.3 (max) Z AI | 67.2 | $1.40 | $4.40 | 85 | 1.88s | 31.2 | |
| 11 | GPT-5.6 Terra (max) OpenAI | 66.7 | $2.00 | $12.00 | 119 | 186.73s | 14.8 | |
| 12 | Claude Opus 5 (Adaptive Reasoning, Medium Effort) Anthropic | 66.5 | $5.00 | $25.00 | 53 | 4.92s | 6.6 | |
| 13 | Gemini 3.7 Flash (high) Google | 66.0 | $0.75 | $3.75 | 323 | 15.15s | 44.0 | |
| 14 | GPT-5.6 Sol (medium) OpenAI | 66.0 | $5.00 | $30.00 | 73 | 3.52s | 5.9 | |
| 15 | Claude Opus 4.8 (Adaptive Reasoning, Max Effort) Anthropic | 65.8 | $5.00 | $25.00 | 56 | 12.05s | 6.6 | |
| 16 | GPT-5.5 (xhigh) OpenAI | 65.6 | $5.00 | $30.00 | 85 | 68.86s | 5.8 | |
| 17 | Qwen3.8 Max Alibaba | 65.0 | $2.00 | $6.00 | 50 | 2.56s | 21.7 | |
| 18 | Qwen3.8 2.4T A95B Alibaba | 64.8 | $2.00 | $6.00 | 45 | 2.53s | 21.6 | |
| 19 | Muse Spark 1.2 (xhigh) Meta | 64.5 | $1.25 | $4.25 | — | — | 32.2 | |
| 20 | Claude Opus 4.7 (Adaptive Reasoning, Max Effort) Anthropic | 64.3 | $5.00 | $25.00 | 48 | 15.61s | 6.4 |
Score per dollar (higher = better value). Only models with pricing data.
Models in the top-right are both fast and capable.
PinchBench evaluates AI models on real-world agent tasks spanning coding, reasoning, tool use, and instruction following. Unlike academic benchmarks that test isolated capabilities, PinchBench combines multiple benchmark dimensions to reflect how models perform as autonomous agents in practical workflows.
PinchBench covers 6 scenarios: Coding Agent (code generation, debugging, terminal use), Reasoning & Logic (math, science, multi-step problems), Instruction Following (format compliance, structured output), Research & Analysis (scientific reasoning, knowledge), Tool Use & Agentic (multi-turn orchestration, planning), and an Overall balanced score.
Each scenario uses a weighted combination of relevant benchmarks. For example, Coding Agent combines LiveCodeBench, TerminalBench, SciCode, and the Artificial Analysis Coding Index. Scores are normalized to 0-100. Cost efficiency is calculated as score divided by price per million tokens.
Academic benchmarks test specific skills in controlled conditions. Real agent tasks require combining multiple skills — a model might score well on individual benchmarks but struggle when tasks require coding + tool use + instruction following simultaneously. PinchBench's weighted scenario scores better approximate this combined performance.
PinchBench data refreshes hourly from the Artificial Analysis API, ensuring you see the latest benchmark scores and pricing for all models.