PinchBench
PinchBench is a public leaderboard benchmarking LLMs on standardized OpenClaw coding tasks, comparing success rate, speed, and cost across 50+ models.
About
PinchBench is a public leaderboard benchmarking LLMs on standardized OpenClaw coding tasks, comparing success rate, speed, and cost across 50+ models. Built by Kilo Code, it runs live benchmarks updated continuously and exposes the data openly — down to individual task scores and run methodology. The leaderboard shows GPT-5.4 and Qwen 3.5 models currently leading, with Claude Sonnet 4.5/4.6 close behind.
AI engineers choosing LLMs for coding agents who want OpenClaw-specific benchmark data rather than general academic benchmarks — particularly useful for cost-performance trade-off decisions at scale.
Pros & Cons
Pros
- check Covers 50+ models with both best and average scores, giving a realistic picture rather than cherry-picked peaks
- check Three-way trade-off view: success rate, speed, and cost in one place for budget-aware model selection
- check Openly reproducible — all tasks and grading criteria are open source, and you can run the benchmark yourself
- check Updated continuously (daily snapshots) so rankings reflect current model versions
- check Includes both official and unofficial runs, with toggles for open-weight models only
Cons
- close Focused exclusively on OpenClaw coding tasks — not a general AI benchmark; results don't transfer to other domains
- close Powered by one company (Kilo Code), which has potential conflicts of interest if Kilo's own models appear in results
- close Success rates vary significantly between best and average runs — a single best-score headline can be misleading
- close Cost metrics depend on current API pricing, which changes frequently
More Infrastructure
Other tools in the same category.