What PinchBench Tests

PinchBench evaluates AI models on 6 real-world agent scenarios: coding agent (LiveCodeBench, TerminalBench, SciCode), reasoning & logic (GPQA, AIME 2025, MATH-500, HLE), instruction following (IFBench, MMLU-Pro), research & analysis, and tool use & agentic (τ²-bench, LCR). For raw academic scores, see our full LLM benchmark leaderboard, and for raw speed numbers, the AI model speed rankings.

What is PinchBench? (AI Agent Benchmark Leaderboard 2026)

PinchBench is a real-world AI agent benchmark that scores 400+ models on production-style agent tasks — coding agent scenarios (LiveCodeBench, TerminalBench, SciCode), reasoning & logic (GPQA Diamond, AIME 2025, HLE, MATH-500), instruction following (IFBench, MMLU-Pro), tool use & agentic (τ²-bench, LCR), and research workflows. Each model receives two composite scores: a Coding Index weighted toward agentic coding performance and an Intelligence Index weighted toward general reasoning.

Live data · Last updated · Refreshed hourly from Artificial Analysis

Top agents today by Coding Index: GPT-5.6 Sol (xhigh) 78.3, Claude Opus 5 (Max) 78.0, GPT-5.6 Sol (max) 77.4. Top agents by Intelligence Index: Claude Opus 5 (Max) II 60.7, Claude Opus 5 (Xhigh) II 60.1, Claude Fable 5 II 59.9. The leaderboard is updated hourly from Artificial Analysis data — no signup required.

How is PinchBench different from GPQA, HLE, or other benchmarks?

Most benchmarks measure single-turn academic performance (GPQA, MMLU-Pro, AIME 2025). PinchBench measures end-to-end agent behavior: can the model write a function, run it, see it fail, and iterate? That requires tool use (τ²-bench, LCR), instruction following (IFBench), and multi-step coding (LiveCodeBench, TerminalBench, SciCode). It's the difference between "can solve a math problem" and "can build a working app from a spec."

Top 3 on PinchBench Today (August 2026)

  1. 🥇Claude Opus 5 (Adaptive Reasoning, Max Effort) — Intelligence Index 60.7, Coding Index 78.0View →
  2. 🥈GPT-5.6 Sol (xhigh) — Coding Index 78.3, Intelligence Index 57.7View →
  3. 🥉Claude Fable 5 — Intelligence Index 59.9, HLE 53.3% #1View →
Live data · Updated hourly

PinchBench — Real-World AI Agent Benchmarks

How do AI models perform on real agent tasks? PinchBench scores 603+ models across coding, reasoning, tool use, and instruction following — with live pricing data.

Models Tested
603
Scenarios
6
Avg Score
22.5
Best Value
Ling 3.0 Flash
Overall

Balanced score across all agent capabilities

intelligence index (15%)coding index (15%)math index (10%)gpqa (10%)livecodebench (10%)ifbench (10%)tau2 (10%)terminalbench hard (10%)hle (10%)
🥇#170.5
Anthropic

Claude Opus 5 (Adaptive Reasoning, Max Effort)

Price
$10.00
Speed
53
Efficiency
7.1
🥈#269.8
Anthropic

Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)

Price
$10.00
Speed
54
Efficiency
7.0
🥉#369.3
Anthropic

Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)

Price
$20.00
Speed
67
Efficiency
3.5
#ModelScoreBarInput $/MOutput $/MSpeedTTFTEfficiency
1
Claude Opus 5 (Adaptive Reasoning, Max Effort)
Anthropic
70.5
$5.00$25.005359.97s7.1
2
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
Anthropic
69.8
$5.00$25.005440.50s7.0
3
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Anthropic
69.3
$10.00$50.0067148.12s3.5
4
GPT-5.6 Sol (max)
OpenAI
69.2
$5.00$30.0065208.71s6.1
5
Claude Opus 5 (Adaptive Reasoning, High Effort)
Anthropic
69.0
$5.00$25.005211.81s6.9
6
Grok 4.6 (high)
SpaceXAI
68.9
$2.00$6.005636.89s23.0
7
GPT-5.6 Sol (xhigh)
OpenAI
68.6
$5.00$30.006350.77s6.1
8
Kimi K3 (max)
Kimi
68.0
$3.00$15.00382.67s11.3
9
GPT-5.6 Sol (high)
OpenAI
67.3
$5.00$30.00709.27s6.0
10
GLM-5.3 (max)
Z AI
67.2
$1.40$4.40851.88s31.2
11
GPT-5.6 Terra (max)
OpenAI
66.7
$2.00$12.00119186.73s14.8
12
Claude Opus 5 (Adaptive Reasoning, Medium Effort)
Anthropic
66.5
$5.00$25.00534.92s6.6
13
Gemini 3.7 Flash (high)
Google
66.0
$0.75$3.7532315.15s44.0
14
GPT-5.6 Sol (medium)
OpenAI
66.0
$5.00$30.00733.52s5.9
15
Claude Opus 4.8 (Adaptive Reasoning, Max Effort)
Anthropic
65.8
$5.00$25.005612.05s6.6
16
GPT-5.5 (xhigh)
OpenAI
65.6
$5.00$30.008568.86s5.8
17
Qwen3.8 Max
Alibaba
65.0
$2.00$6.00502.56s21.7
18
Qwen3.8 2.4T A95B
Alibaba
64.8
$2.00$6.00452.53s21.6
19
Muse Spark 1.2 (xhigh)
Meta
64.5
$1.25$4.2532.2
20
Claude Opus 4.7 (Adaptive Reasoning, Max Effort)
Anthropic
64.3
$5.00$25.004815.61s6.4

💰 Best Cost Efficiency — Overall

Score per dollar (higher = better value). Only models with pricing data.

1
Ling 3.0 Flash
409.3$0.11
2
Qwen3.5 4B (Reasoning)
358.3$0.06
3
Hy3-preview (Reasoning)
351.0$0.10
4
Qwen3.5 4B (Non-reasoning)
303.3$0.06
5
DeepSeek V4 Flash (Reasoning, Max Effort)
292.6$0.17
6
Hy3-preview (Non-reasoning)
271.4$0.10
7
MiMo-V2.5
270.9$0.17
8
DeepSeek V4 Flash (Reasoning, High Effort)
270.8$0.17
9
Gemma 4 E4B (Reasoning)
270.0$0.04
10
DeepSeek V4 Flash (Non-reasoning)
248.3$0.12

⚡ Score vs Speed — Overall

Models in the top-right are both fast and capable.

Celeris
Celeris-1
Score
13.4
Speed
1484
Inception
Mercury 2
Score
26.5
Speed
835
Google
Gemini 3.7 Flash (high)
Score
66.0
Speed
323
Google
Gemini 3.7 Flash (medium)
Score
62.5
Speed
317
Google
Gemini 3.7 Flash (low)
Score
61.0
Speed
303
Google
Gemini 3.5 Flash-Lite
Score
43.4
Speed
332
Google
Gemini 3.1 Flash-Lite
Score
30.2
Speed
346
Google
Gemini 3.6 Flash (high)
Score
60.4
Speed
222
NVIDIA
Nemotron 3.5 Lightning
Score
25.2
Speed
302
NVIDIA
Nemotron 3 Nano Omni 30B A3B Reasoning
Score
14.4
Speed
326

Frequently Asked Questions

What is PinchBench and how does it differ from traditional benchmarks?

PinchBench evaluates AI models on real-world agent tasks spanning coding, reasoning, tool use, and instruction following. Unlike academic benchmarks that test isolated capabilities, PinchBench combines multiple benchmark dimensions to reflect how models perform as autonomous agents in practical workflows.

Which scenarios does PinchBench test?

PinchBench covers 6 scenarios: Coding Agent (code generation, debugging, terminal use), Reasoning & Logic (math, science, multi-step problems), Instruction Following (format compliance, structured output), Research & Analysis (scientific reasoning, knowledge), Tool Use & Agentic (multi-turn orchestration, planning), and an Overall balanced score.

How are scores calculated?

Each scenario uses a weighted combination of relevant benchmarks. For example, Coding Agent combines LiveCodeBench, TerminalBench, SciCode, and the Artificial Analysis Coding Index. Scores are normalized to 0-100. Cost efficiency is calculated as score divided by price per million tokens.

Why do real-world results differ from academic benchmarks?

Academic benchmarks test specific skills in controlled conditions. Real agent tasks require combining multiple skills — a model might score well on individual benchmarks but struggle when tasks require coding + tool use + instruction following simultaneously. PinchBench's weighted scenario scores better approximate this combined performance.

How often is the data updated?

PinchBench data refreshes hourly from the Artificial Analysis API, ensuring you see the latest benchmark scores and pricing for all models.