Get LLM-specific scores for discoverability, tool calls, and errors. Real sandbox execution across Claude Code, Cursor, Codex and 11 agents, not just prompt analysis.
Benchmark your API, MCP, or CLI across all major coding agents
Track how your API performs with AI coding agents
Which coding agents work best with your API
| Agent / Average | Payments | Billing | Checkout | Connect | Webhooks |
|---|---|---|---|---|---|
1 | 97% | ||||
2 | 91% | ||||
3 | 82% | ||||
4 | 78% | ||||
5 | 71% | ||||
6 | 66% |
How each endpoint performs across eval types
/v1/payment_intents/v1/chargesConsumer AI visibility tools track whether ChatGPT mentions your brand. But for APIs, MCP servers, and CLIs, a mention means nothing if the generated code doesn't compile.
When Claude Code recommends your tool but the code fails, developers switch to a competitor. You need to measure what actually matters: can agents discover, call, and use your endpoints and tools?
Agent Usability benchmarks your API, MCP, or CLI with real sandbox execution. Get scores for discoverability, tool calls, and error rates broken down by agent and endpoint.
Every endpoint. Every agent. Every metric.
Spin up actual environments and run your API, MCP server, or CLI against real coding agents. This is actual code generation and execution with pass/fail results, not prompt analysis.
Write your own eval prompts. Test specific endpoints, MCP tools, CLI commands, auth flows, and edge cases. Define what success looks like.
See exactly where agents fail: 404s, auth errors, malformed requests, missing docs. Get the specific URLs and step-by-step traces to debug.
Get separate scores for discoverability, tool calls, and error rates. Each endpoint, agent, and metric is broken down so you know what to fix.
See how your API, MCP, or CLI ranks against competitors on Devtool Arena. Track your position over time and identify gaps.
Claude Code, Cursor, Codex, GitHub Copilot, Gemini CLI, and more. Each agent behaves differently, so know your compatibility scores across all of them.
See how your API ranks against competitors on our free public leaderboard. Based on real benchmarks across discoverability, tool calls, and error rates.
View the leaderboardFind out how well AI coding agents can use your API. Get specific recommendations to improve.