Advertisement

What PinchBench Tests

PinchBench evaluates AI models on 6 real-world agent scenarios: coding agent (LiveCodeBench, TerminalBench, SciCode), reasoning & logic (GPQA, AIME 2025, MATH-500, HLE), instruction following (IFBench, MMLU-Pro), research & analysis, and tool use & agentic (τ²-bench, LCR). For raw academic scores, see our full LLM benchmark leaderboard, and for raw speed numbers, the AI model speed rankings.

What is PinchBench? (AI Agent Benchmark Leaderboard 2026)

PinchBench is a real-world AI agent benchmark that scores 400+ models on production-style agent tasks — coding agent scenarios (LiveCodeBench, TerminalBench, SciCode), reasoning & logic (GPQA Diamond, AIME 2025, HLE, MATH-500), instruction following (IFBench, MMLU-Pro), tool use & agentic (τ²-bench, LCR), and research workflows. Each model receives two composite scores: a Coding Index weighted toward agentic coding performance and an Intelligence Index weighted toward general reasoning.

Live data · Last updated · Refreshed hourly from Artificial Analysis

Top agents today by Coding Index: GPT-5.6 Sol (xhigh) 78.3, Claude Opus 5 (Max) 78.0, GPT-5.6 Sol (max) 77.4. Top agents by Intelligence Index: Claude Opus 5 (Max) II 60.7, Claude Opus 5 (Xhigh) II 60.1, Claude Fable 5 II 59.9. The leaderboard is updated hourly from Artificial Analysis data — no signup required.

How is PinchBench different from GPQA, HLE, or other benchmarks?

Most benchmarks measure single-turn academic performance (GPQA, MMLU-Pro, AIME 2025). PinchBench measures end-to-end agent behavior: can the model write a function, run it, see it fail, and iterate? That requires tool use (τ²-bench, LCR), instruction following (IFBench), and multi-step coding (LiveCodeBench, TerminalBench, SciCode). It's the difference between "can solve a math problem" and "can build a working app from a spec."

Top 3 on PinchBench Today (August 2026)

  1. 🥇Claude Opus 5 (Adaptive Reasoning, Max Effort) — Intelligence Index 60.7, Coding Index 78.0View →
  2. 🥈GPT-5.6 Sol (xhigh) — Coding Index 78.3, Intelligence Index 57.7View →
  3. 🥉Claude Fable 5 — Intelligence Index 59.9, HLE 53.3% #1View →
Live data · Updated hourly

PinchBench — Real-World AI Agent Benchmarks

How do AI models perform on real agent tasks? PinchBench scores 641+ models across coding, reasoning, tool use, and instruction following — with live pricing data.

Models Tested
641
Scenarios
6
Avg Score
20.1
Best Value
Agnes 3.0 Flash
⭐ Overall

Balanced score across all agent capabilities

intelligence index (15%)coding index (15%)math index (10%)gpqa (10%)livecodebench (10%)ifbench (10%)tau2 (10%)terminalbench hard (10%)hle (10%)
🥇#167.5
Anthropic

Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)

Price
$20.00
Speed
65
Efficiency
3.4
🥈#267.0
Anthropic

Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)

Price
$20.00
Speed
60
Efficiency
3.3
🥉#365.1
Anthropic

Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)

Price
$20.00
Speed
55
Efficiency
3.3
#ModelScoreBarInput $/MOutput $/MSpeedTTFTEfficiency
1
Claude Fable 5.1 (Adaptive Reasoning, Max Effort, Default Fallback)
Anthropic
67.5
$10.00$50.0065210.79s3.4
2
Claude Fable 5.1 (Adaptive Reasoning, Xhigh Effort, Default Fallback)
Anthropic
67.0
$10.00$50.006088.63s3.3
3
Claude Fable 5.1 (Adaptive Reasoning, High Effort, Default Fallback)
Anthropic
65.1
$10.00$50.005514.87s3.3
4
GPT-6 Astra (max)
OpenAI
64.9
$10.00$50.0060289.72s3.2
5
Claude Opus 5 (Adaptive Reasoning, Max Effort)
Anthropic
64.4
$5.00$25.005248.08s6.4
6
GPT-6 Astra (xhigh)
OpenAI
64.2
$10.00$50.0055125.02s3.2
7
GPT-6 Astra (high)
OpenAI
64.0
$10.00$50.005336.95s3.2
8
Claude Opus 5 (Adaptive Reasoning, Xhigh Effort)
Anthropic
63.4
$5.00$25.005326.77s6.3
9
GPT-6 Astra (medium)
OpenAI
63.2
$10.00$50.00523.85s3.2
10
Claude Fable 5 (Adaptive Reasoning, Max Effort, Opus 4.8 Fallback)
Anthropic
63.1
$10.00$50.006188.23s3.2
11
Claude Fable 5.1 (Adaptive Reasoning, Medium Effort, Default Fallback)
Anthropic
63.1
$10.00$50.00548.69s3.2
12
Claude Opus 5 (Adaptive Reasoning, High Effort)
Anthropic
62.4
$5.00$25.005111.51s6.2
13
GPT-5.6 Sol (max)
OpenAI
62.3
$4.00$20.0063119.96s7.8
14
Muse Spark 1.3 (max)
Meta
62.0
$1.25$4.2523323.48s31.0
15
GPT-5.6 Sol (xhigh)
OpenAI
61.2
$4.00$20.006038.98s7.6
16
Claude Fable 5.1 (Adaptive Reasoning, Low Effort, Default Fallback)
Anthropic
61.1
$10.00$50.00534.71s3.1
17
GPT-6 Astra (low)
OpenAI
60.9
$10.00$50.00512.59s3.0
18
Muse Spark 1.3 (xhigh)
Meta
60.9
$1.25$4.2521722.70s30.4
19
Grok 4.6 (high)
SpaceXAI
60.6
$2.00$6.005940.86s20.2
20
Grok 4.6 (xhigh)
SpaceXAI
60.1
$2.00$6.005541.74s20.0

💰 Best Cost Efficiency — Overall

Score per dollar (higher = better value). Only models with pricing data.

1
Agnes 3.0 Flash
473.3$0.07
2
Ling 3.0 Flash
329.6$0.11
3
Agnes 2.5 Pro Beta
325.0$0.15
4
Qwen3.5 4B (Reasoning)
297.5$0.06
5
HyperNova 60B 2605 (high, based on gpt-oss-120b)
268.5$0.07
6
Qwen3.5 4B (Non-reasoning)
259.2$0.06
7
Granite 4.2 3B
250.9$0.05
8
Qwen3.8-Flash-Next
245.7$0.23
9
DeepSeek V4 Flash (Reasoning, Max Effort)
240.5$0.17
10
GLM-5.3-Flash
238.2$0.24

⚡ Score vs Speed — Overall

Models in the top-right are both fast and capable.

Celeris
Celeris-1
Score
10.4
Speed
1460
Inception
Mercury 2
Score
21.3
Speed
881
Google
Gemini 3.5 Flash-Lite
Score
36.0
Speed
365
Google
Gemini 3.8 Flash (high)
Score
58.8
Speed
303
Google
Gemini 3.7 Flash (low)
Score
53.9
Speed
308
Google
Gemini 2.5 Flash-Lite (Reasoning)
Score
8.5
Speed
416
Google
Gemini 3.7 Flash (high)
Score
57.7
Speed
288
Google
Gemini 3.7 Flash (medium)
Score
55.5
Speed
289
InclusionAI
Ling 3.0 Flash
Score
35.6
Speed
323
Multiverse Computing
HyperNova 60B 2605 (high, based on gpt-oss-120b)
Score
17.4
Speed
353

Frequently Asked Questions

What is PinchBench and how does it differ from traditional benchmarks?

PinchBench evaluates AI models on real-world agent tasks spanning coding, reasoning, tool use, and instruction following. Unlike academic benchmarks that test isolated capabilities, PinchBench combines multiple benchmark dimensions to reflect how models perform as autonomous agents in practical workflows.

Which scenarios does PinchBench test?

PinchBench covers 6 scenarios: Coding Agent (code generation, debugging, terminal use), Reasoning & Logic (math, science, multi-step problems), Instruction Following (format compliance, structured output), Research & Analysis (scientific reasoning, knowledge), Tool Use & Agentic (multi-turn orchestration, planning), and an Overall balanced score.

How are scores calculated?

Each scenario uses a weighted combination of relevant benchmarks. For example, Coding Agent combines LiveCodeBench, TerminalBench, SciCode, and the Artificial Analysis Coding Index. Scores are normalized to 0-100. Cost efficiency is calculated as score divided by price per million tokens.

Why do real-world results differ from academic benchmarks?

Academic benchmarks test specific skills in controlled conditions. Real agent tasks require combining multiple skills — a model might score well on individual benchmarks but struggle when tasks require coding + tool use + instruction following simultaneously. PinchBench's weighted scenario scores better approximate this combined performance.

How often is the data updated?

PinchBench data refreshes hourly from the Artificial Analysis API, ensuring you see the latest benchmark scores and pricing for all models.