AI Coding Agent Leaderboard
PinchBench success rates for AI coding agents, cross-referenced with ModelPriceLab stored listed-price records so you can compare benchmark performance and public prices.
Success Rate Leaderboard
Success rate = percentage of standardized OpenClaw agent tasks completed successfully. Graded via automated checks and LLM judge.
Benchmark data from PinchBench (pinchbench.com), refreshed daily; latest run 2026-10-05. Price data from ModelPriceLab. For informational purposes only.
About PinchBench
PinchBench is an open-source benchmarking system that evaluates LLMs as OpenClaw coding agents across 15 standardized Code & DevOps tasks. Unlike traditional LLM benchmarks, PinchBench focuses on tool usage, multi-step reasoning, handling ambiguous instructions, and practical outcomes.
23 Benchmark Tasks
Covering calendar creation, code writing, document summarization, email triage, market research, and more. Each task is graded via automated checks, LLM judge, or a hybrid approach.
Key Takeaways
Two Perfect Scores
Claude Opus 4.6 and Inception Mercury 2 tie at 100% success; Claude Opus 4.8 Fast follows at 94.5%.
Anthropic Owns the Top
All 6 Claude models are ranked; 5 of the top 13 spots are Anthropic, from Opus 4.6 (100%) to Haiku 4.5 (90.4%).
57x Cost Spread
DeepSeek-V4 Flash scores 91.5% at $1.44 per run, while top-tier models cost $47–$165. Compare prices before assuming a premium.
Better model decisions come from seeing benchmark scores and API pricing together
Explore the full price matrix, scenario leaderboards, and solution library on ModelPriceLab to find the optimal combination for your use case.