Skip to content
ModelPriceLab

Search model prices

Find source-linked rates by model, vendor, or platform.

← Back to home
Benchmark · PinchBench

AI Coding Agent Leaderboard

PinchBench success rates for AI coding agents, cross-referenced with ModelPriceLab stored listed-price records so you can compare benchmark performance and public prices.

57 models· 15 Code & DevOps tasks· Price overlay·

Success Rate Leaderboard

Success rate = percentage of standardized OpenClaw agent tasks completed successfully. Graded via automated checks and LLM judge.

#1
🦞Xiaomi MiMo-V2.5 Pro
xiaomi/mimo-v2.5-pro
Success
98.0%
Run Cost
11.39
Input / 1M
$0.43
Output / 1M
$0.87
#2
🦀Xiaomi MiMo-V2.5
xiaomi/mimo-v2.5
Success
97.9%
Run Cost
5.44
Input / 1M
$0.14
Output / 1M
$0.28
#3
🦐MiniMax M2.7
minimax/minimax-m2.7
Success
96.2%
Run Cost
2.08
Input / 1M
$0.21
Output / 1M
$0.84
#4
Claude Opus 4.7
anthropic/claude-opus-4.7
Success
95.9%
Run Cost
38.04
Input / 1M
$2.50
Output / 1M
$12.50
#5
ByteDance Seed 2.0 Lite
bytedance-seed/seed-2.0-lite
Success
94.9%
Run Cost
1.74
Input / 1M
$0.25
Output / 1M
$2.00
#6
GLM 5V Turbo
z-ai/glm-5v-turbo
Success
94.9%
Run Cost
12.81
Input / 1M
$1.20
Output / 1M
$4.00
#7
Claude Opus 4.8 Fast
anthropic/claude-opus-4.8-fast
Success
94.5%
Run Cost
159.60
Input / 1M
$2.50
Output / 1M
$12.50
#8
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Success
94.5%
Run Cost
15.18
Input / 1M
$1.50
Output / 1M
$7.50
#9
DeepSeek-V4 Flash
deepseek/deepseek-v4-flash
Success
94.0%
Run Cost
0.88
Input / 1M
$0.15
Output / 1M
$0.60
#10
Gemma 4 26B
google/gemma-4-26b-a4b-it
Success
93.9%
Run Cost
0.45
Input / 1M
$0.07
Output / 1M
$0.23
#11
Qwen3.7 Max
qwen/qwen3.7-max
Success
93.8%
Run Cost
19.43
Input / 1M
$1.48
Output / 1M
$4.42
#12
Grok 4.3
x-ai/grok-4.3
Success
93.7%
Run Cost
13.00
Input / 1M
$1.25
Output / 1M
$2.50
#13
Gemma 4 31B
google/gemma-4-31b-it
Success
93.6%
Run Cost
1.03
Input / 1M
$0.09
Output / 1M
$0.34
#14
Qwen3.6 Plus
qwen/qwen3.6-plus
Success
93.6%
Run Cost
6.97
Input / 1M
$0.33
Output / 1M
$1.95
#15
Nemotron 3.5 Lightning 30B
nvidia/nemotron-3.5-lightning-30b-a3b
Success
93.4%
Run Cost
-
Input / 1M
$0
Output / 1M
$0
#16
Mistral Devstral 2512
mistralai/devstral-2512
Success
93.4%
Run Cost
2.30
Input / 1M
$0.40
Output / 1M
$2.00
#17
Claude Opus 4.8
anthropic/claude-opus-4.8
Success
92.8%
Run Cost
80.10
Input / 1M
$2.50
Output / 1M
$12.50
#18
GLM 5.1
z-ai/glm-5.1
Success
92.6%
Run Cost
-
Input / 1M
$1.40
Output / 1M
$4.40
#19
DeepSeek-V4 Pro
deepseek/deepseek-v4-pro
Success
92.3%
Run Cost
-
Input / 1M
$0.66
Output / 1M
$1.98
#20
GPT-5.5
openai/gpt-5.5
Success
92.3%
Run Cost
22.17
Input / 1M
$2.50
Output / 1M
$15.00
#21
Gemini 3 Flash Preview
google/gemini-3-flash-preview
Success
92.3%
Run Cost
3.09
Input / 1M
$0.50
Output / 1M
$3.00
#22
Nemotron 3 Ultra 550B
nvidia/nemotron-3-ultra-550b-a55b
Success
92.0%
Run Cost
-
Input / 1M
$0.50
Output / 1M
$2.20
#23
Gemini 3.1 Flash Lite
google/gemini-3.1-flash-lite
Success
91.3%
Run Cost
3.79
Input / 1M
$0.25
Output / 1M
$1.50
#24
GPT-5.6 Sol
openai/gpt-5.6-sol
Success
90.9%
Run Cost
65.60
Input / 1M
$4.00
Output / 1M
$20.00
#25
Qwen3.6 Flash
qwen/qwen3.6-flash
Success
89.9%
Run Cost
12.59
Input / 1M
$0.19
Output / 1M
$1.13
#26
Claude Haiku 4.5
anthropic/claude-haiku-4.5
Success
89.8%
Run Cost
1.69
Input / 1M
$1.00
Output / 1M
$5.00
#27
Gemini 3.5 Flash
google/gemini-3.5-flash
Success
89.8%
Run Cost
25.84
Input / 1M
$1.50
Output / 1M
$9.00
#28
Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview
Success
89.6%
Run Cost
15.84
Input / 1M
$2.00
Output / 1M
$12.00
#29
Grok 4.5
x-ai/grok-4.5
Success
89.5%
Run Cost
6.15
Input / 1M
$2.00
Output / 1M
$6.00
#30
Grok 4.20
x-ai/grok-4.20
Success
89.4%
Run Cost
16.68
Input / 1M
$1.25
Output / 1M
$2.50
#31
Step 3.5 Flash
stepfun/step-3.5-flash
Success
89.3%
Run Cost
0.50
Input / 1M
$0.10
Output / 1M
$0.30
#32
Sakana Fugu Ultra
sakana/fugu-ultra
Success
89.0%
Run Cost
86.11
Input / 1M
$5.00
Output / 1M
$30.00
#33
GLM 5 Turbo
z-ai/glm-5-turbo
Success
88.9%
Run Cost
5.89
Input / 1M
$1.20
Output / 1M
$4.00
#34
GPT-5.4
openai/gpt-5.4
Success
88.3%
Run Cost
10.38
Input / 1M
$1.25
Output / 1M
$7.50
#35
Kimi K2.7 Code
moonshotai/kimi-k2.7-code
Success
88.1%
Run Cost
10.32
Input / 1M
$0.67
Output / 1M
$3.35
#36
GPT-5.4 Mini
openai/gpt-5.4-mini
Success
87.8%
Run Cost
3.40
Input / 1M
$0.38
Output / 1M
$2.25
#37
Trinity Large Thinking
arcee-ai/trinity-large-thinking
Success
87.0%
Run Cost
1.29
Input / 1M
$0.25
Output / 1M
$0.80
#38
GPT-5.6 Luna
openai/gpt-5.6-luna
Success
86.7%
Run Cost
14.36
Input / 1M
$0.10
Output / 1M
$0.60
#39
Aion 3.0
aion-labs/aion-3.0
Success
85.8%
Run Cost
30.26
Input / 1M
$3.00
Output / 1M
$6.00
#40
Grok Build 0.1
x-ai/grok-build-0.1
Success
85.8%
Run Cost
19.03
Input / 1M
$1.00
Output / 1M
$2.00
#41
Kimi K2.5
moonshotai/kimi-k2.5
Success
83.4%
Run Cost
2.77
Input / 1M
$0.45
Output / 1M
$2.25
#42
Mistral Small 2603
mistralai/mistral-small-2603
Success
82.8%
Run Cost
2.17
Input / 1M
$0.15
Output / 1M
$0.60
#43
GLM 5.2
z-ai/glm-5.2
Success
81.9%
Run Cost
18.20
Input / 1M
$0.02
Output / 1M
$16.00
#44
Mistral Large 2512
mistralai/mistral-large-2512
Success
81.8%
Run Cost
0.48
Input / 1M
$0.50
Output / 1M
$1.50
#45
Ling 2.6 1T
inclusionai/ling-2.6-1t
Success
81.4%
Run Cost
7.00
Input / 1M
$0.02
Output / 1M
$0.06
#46
GPT-5.4 Nano
openai/gpt-5.4-nano
Success
78.6%
Run Cost
0.78
Input / 1M
$0.20
Output / 1M
$1.25
#47
GPT-5.6 Terra
openai/gpt-5.6-terra
Success
76.0%
Run Cost
24.13
Input / 1M
$1.00
Output / 1M
$6.00
#48
Trinity Large Preview
arcee-ai/trinity-large-preview
Success
72.2%
Run Cost
1.90
Input / 1M
$0.25
Output / 1M
$0.80
#49
Nemotron 3 Super 120B
nvidia/nemotron-3-super-120b-a12b
Success
67.0%
Run Cost
-
Input / 1M
$0.08
Output / 1M
$0.45
#50
GPT OSS 120B
openai/gpt-oss-120b
Success
64.7%
Run Cost
0.30
Input / 1M
$0.04
Output / 1M
$0.17
#51
Amazon Nova 2 Lite V1
amazon/nova-2-lite-v1
Success
50.8%
Run Cost
0.84
Input / 1M
$0.30
Output / 1M
$2.50
#52
GPT OSS 20B
openai/gpt-oss-20b
Success
50.0%
Run Cost
0.31
Input / 1M
$0.02
Output / 1M
$0.09
#53
GPT-5.5 Pro
openai/gpt-5.5-pro
Success
48.9%
Run Cost
153.44
Input / 1M
$15.00
Output / 1M
$90.00
#54
Claude Fable 5
anthropic/claude-fable-5
Success
37.8%
Run Cost
114.16
Input / 1M
$10.00
Output / 1M
$50.00
#55
Claude Sonnet 4
anthropic/claude-sonnet-4
Success
22.3%
Run Cost
21.19
Input / 1M
$3.00
Output / 1M
$15.00
#56
Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instruct
Success
19.0%
Run Cost
5.00
Input / 1M
$0.40
Output / 1M
$0.40
#57
Llama 4 Scout
meta-llama/llama-4-scout
Success
8.4%
Run Cost
0.24
Input / 1M
$0.10
Output / 1M
$0.30

Benchmark data from PinchBench (pinchbench.com), refreshed daily; latest run 2026-10-05. Price data from ModelPriceLab. For informational purposes only.

About PinchBench

PinchBench is an open-source benchmarking system that evaluates LLMs as OpenClaw coding agents across 15 standardized Code & DevOps tasks. Unlike traditional LLM benchmarks, PinchBench focuses on tool usage, multi-step reasoning, handling ambiguous instructions, and practical outcomes.

Tool Usage
Can the model call tools with correct parameters
Multi-step Reasoning
Can it chain actions to complete complex tasks
Real-world Messiness
Can it handle ambiguous and incomplete information
Practical Outcomes
Did it actually create the file, send the email

23 Benchmark Tasks

Covering calendar creation, code writing, document summarization, email triage, market research, and more. Each task is graded via automated checks, LLM judge, or a hybrid approach.

✅Sanity Check
Automated
📅Calendar Event
Automated
📈Stock Research
Automated
✍️Blog Writing
LLM Judge
🌤️Weather Script
Automated
📄Doc Summary
LLM Judge
🎤Conference Research
LLM Judge
✉️Email Drafting
LLM Judge
🧠Memory Retrieval
Automated
📁File Structure
Automated
🔄API Workflow
Hybrid
🔌Skill Install
Automated
🔍Skill Search
Automated
🎨Image Generation
Hybrid
🤖Humanize AI Text
LLM Judge
📊Research Summary
LLM Judge
📬Email Triage
Hybrid
🔎Email Search
Hybrid
🏢Market Research
Hybrid
📑CSV/Excel Analysis
Hybrid
👶ELI5 PDF Summary
LLM Judge
📖Report Comprehension
Automated
💾Knowledge Persistence
Hybrid

Key Takeaways

Two Perfect Scores

Claude Opus 4.6 and Inception Mercury 2 tie at 100% success; Claude Opus 4.8 Fast follows at 94.5%.

Anthropic Owns the Top

All 6 Claude models are ranked; 5 of the top 13 spots are Anthropic, from Opus 4.6 (100%) to Haiku 4.5 (90.4%).

57x Cost Spread

DeepSeek-V4 Flash scores 91.5% at $1.44 per run, while top-tier models cost $47–$165. Compare prices before assuming a premium.

Better model decisions come from seeing benchmark scores and API pricing together

Explore the full price matrix, scenario leaderboards, and solution library on ModelPriceLab to find the optimal combination for your use case.