NewIntroducing Agent Payments Arena
Lightsage

Benchmark your API, MCP, or CLI across coding agents

Get LLM-specific scores for discoverability, tool calls, and errors. Real sandbox execution across Claude Code, Cursor, Codex and 11 agents, not just prompt analysis.

Benchmark your API, MCP, or CLI across all major coding agents

Claude CodeClaude Code
CursorCursor
CodexCodex
OpenCodeOpenCode
HermesHermes
+6 more
API Performance

Track how your API performs with AI coding agents

Run Evals
Coding Agent Leaderboard

Which coding agents work best with your API

Agent / AveragePayments
1
Claude Code
Claude Code94%
97%
2
Codex
Codex87%
91%
3
Cursor
Cursor76%
82%
4
GitHub Copilot
GitHub Copilot71%
78%
5
Hermes
Hermes63%
71%
6
OpenCode
OpenCode58%
66%
Endpoint Evaluations

How each endpoint performs across eval types

Payments API(5 endpoints)
82% pass rate
POST/v1/payment_intents
94%
POST/v1/charges
88%
11
coding agents benchmarked
3
score dimensions
Every
endpoint evaluated

Visibility ≠ Usability

Consumer AI visibility tools track whether ChatGPT mentions your brand. But for APIs, MCP servers, and CLIs, a mention means nothing if the generated code doesn't compile.

When Claude Code recommends your tool but the code fails, developers switch to a competitor. You need to measure what actually matters: can agents discover, call, and use your endpoints and tools?

Agent Usability benchmarks your API, MCP, or CLI with real sandbox execution. Get scores for discoverability, tool calls, and error rates broken down by agent and endpoint.

Endpoint-level benchmarks

Every endpoint. Every agent. Every metric.

Real Sandbox Execution

Spin up actual environments and run your API, MCP server, or CLI against real coding agents. This is actual code generation and execution with pass/fail results, not prompt analysis.

Custom Eval Scenarios

Write your own eval prompts. Test specific endpoints, MCP tools, CLI commands, auth flows, and edge cases. Define what success looks like.

Error & Failure Detection

See exactly where agents fail: 404s, auth errors, malformed requests, missing docs. Get the specific URLs and step-by-step traces to debug.

LLM-Specific Scores

Get separate scores for discoverability, tool calls, and error rates. Each endpoint, agent, and metric is broken down so you know what to fix.

Competitor Benchmarks

See how your API, MCP, or CLI ranks against competitors on Devtool Arena. Track your position over time and identify gaps.

11 Coding Agents

Claude Code, Cursor, Codex, GitHub Copilot, Gemini CLI, and more. Each agent behaves differently, so know your compatibility scores across all of them.

Free & Public

Devtool Arena

See how your API ranks against competitors on our free public leaderboard. Based on real benchmarks across discoverability, tool calls, and error rates.

View the leaderboard
#1StripeStripe
94%
#2TwilioTwilio
91%
#3OpenAIOpenAI
87%

Is your API agent-ready?

Find out how well AI coding agents can use your API. Get specific recommendations to improve.