API Agent Landscape
A market map for API teams: which products agents can discover, understand, and use in real tasks.
API Companies
174
344 scored API runs
Categories Mapped
41
API markets with benchmark signal
Average API Score
57
Mean task-completion readiness
Avg Run Cost
$1.30
3m 54s average runtime
Hardest API Category
9.8 error outputs per run
Market Position
Full landscape context across API, CLI, and MCP. X axis: readiness and discovery. Y axis: observed eval success.
Landscape leaders
The strongest companies across the landscape.
Overall ranks companies that are scored by both agents on API, CLI, and MCP. The interface leaderboards show the best companies within each entry point.
Both agents across API, CLI, MCP
Overall leaders
By interface
API
Claude Code + Codex
Storage
Inference
Payment
Voice TTS
Observability
Geocoding
CLI
Claude Code + Codex
Voice Infra
Auth
Browser
Cloud Hosting
Sandboxes
Sandboxes
MCP
Claude Code + Codex
Search
Voice TTS
Search
Cloud Hosting
Payment
Auth
Agent Experience
How does an agent use your API?
Agent experience means the concrete path an agent takes to use an API: find the source of truth, read the right reference, authenticate, make calls, handle failures, and verify the result.
The trace tells us where that path gets harder. Search calls show discovery work, docs calls show reference usage, auth calls show setup friction, and error outputs show recovery work.
All agents API usage path
API runs only · 274 benchmark attempts
Avg score
57Discover the API surface
The agent chooses between official docs, search results, examples, generated context, CLI, and MCP surfaces.
Load the right reference
66% of API runs touched docs. When docs were touched, average score moved +33 points higher.
Authenticate the request
Keys, tokens, headers, scopes, env vars, and account setup determine whether the first real call can run.
Call endpoints and use SDKs
A typical run takes 3m 35s; slow execution increases the cost of agent-driven API usage.
Handle errors and verify output
3.5 reported errors per run on average. This captures retries, failed calls, and result checking.
Agent-readable lift
llms.txt +23 · OpenAPI +11
Operating cost
3m 35s · $1.03 per run
Hardest API category
Stablecoin · 35
Successful automation
75% pass rate in this view
Agent Surface Map
Each cell shows the average final score for that surface-agent slice. Top company is the highest-scoring company in the slice.
Claude Code
Coding agent
Codex
Coding agent
API
API surface
Avg final score
Top company
Companies scored
174
Categories
41
Avg final score
Top company
Companies scored
170
Categories
40
CLI
CLI surface
Avg final score
Top company
Companies scored
95
Categories
22
Avg final score
Top company
Companies scored
60
Categories
20
MCP
MCP surface
Avg final score
Top company
Companies scored
79
Categories
25
Avg final score
Top company
Companies scored
77
Categories
25
API Usage Readout
Plain-English usage signals for API teams
174 API companies are benchmarked
344 scored API runs test whether agents can find docs, authenticate, call endpoints, and verify results.
Average API task score is 57
Lower scores usually mean the agent got stuck in discovery, auth setup, endpoint execution, or recovery.
Browser has the clearest API path
Agents completed these tasks most consistently: 4 companies average 76 across 8 API runs.
Boxis easiest for agents to use right now
Its Storage tasks average 89 across 2 API runs.
Agent Automation depends on the agent
The leading agent is ahead by 21 points. Scores are 69 vs 48, so the same API tasks are not equally easy for both agents.
API usage audit
The important question is whether the agent can complete the API task.
These findings are computed from 274 API runs and 9,115 normalized tool calls. They summarize the agent API workflow: how much searching happens, whether docs are actually used, how much auth/setup work appears, and where endpoint calls fall into recovery loops.
Usage baseline
Most API tasks still require substantial agent effort
274 API runs averaged 3m 35s, 16 tool calls, and $1.03 per attempt.
Docs moment
When agents reach docs, completion gets much easier
Docs-used runs scored 68 on average versus 35 when docs were not touched.
Context signal
Docs used in trace gives agents a clearer starting point
Runs with this signal averaged 68, compared with 35 without it.
Recovery tax
Stablecoin
Agents struggle here because tasks complete less reliably, runs produce 5.4 error outputs on average, agents need 23.3 tool calls per run, and completion takes 6m 3s on average. Example company: BlindPay with 14 errors.
Where agents struggle
These API categories create the most execution friction.
A category ranks here when agents score lower while spending more time, tool calls, error recovery, or cost. Use this to find where docs, auth examples, SDK setup, or endpoint behavior may be slowing agent users down.
Stablecoin
Low completionRecovery loopsMany tool callsSlow runsAvg task score
35Runtime
Tool calls
Errors
Cost
Cloud Hosting
Low completionRecovery loopsMany tool callsSlow runsAvg task score
45Runtime
Tool calls
Errors
Cost
Voice Telephony
Low completionRecovery loopsMany tool callsSlow runsAvg task score
47Runtime
Tool calls
Errors
Cost
Avg task score
67Runtime
Tool calls
Errors
Cost
Voice Infra
Recovery loopsMany tool callsSlow runsAvg task score
54Runtime
Tool calls
Errors
Cost
Auth
Recovery loopsMany tool callsSlow runsAvg task score
69Runtime
Tool calls
Errors
Cost
Video Agent
Recovery loopsMany tool callsSlow runsAvg task score
62Runtime
Tool calls
Errors
Cost
Tooling distribution
Agent tooling is uneven across categories.
Some categories have broad API coverage but few CLI or MCP options. Others give agents multiple ways to complete the same task. The distribution shows which companies have more than one agent-facing surface.
Sandboxes
E2Bleads at 78
Companies
10
Code Review
Macroscopeleads at 17
Companies
5
Voice Telephony
Telnyxleads at 74
Companies
9
Voice STT
Deepgramleads at 71
Companies
6
Durable Workflow
Temporalleads at 78
Companies
8
Inference
Fireworks AIleads at 86
Companies
10
Stablecoin
Circleleads at 59
Companies
7
Voice TTS
ElevenLabsleads at 83
Companies
6
E-Signature
Docusignleads at 74
Companies
3
Category leaders
Each agent has its own category map.
Current category leaders split by agent and surface.
Claude leaders
Claude Code category leaders
Codex leaders
Codex category leaders
Matched surface comparison
CLI vs MCP
This comparison uses 98 matched provider-agent pairs across 61 companies and 19 categories. A pair only counts when the same company has both a CLI and MCP benchmark row for the same coding agent. In this matched set, CLI completes more often, CLI is more runtime-efficient, and CLI produces fewer errors.
Interpretation
CLI is the safer default for reliability-sensitive workflows.
CLI finishes more often and runs faster on the matched set. It has a +12 pts success-rate advantage and saves -2m 9s per run versus MCP.
MCP makes the stronger structured-interface case once the task completes: higher average score, fewer tool calls, and fewer errors per run. Its reported run cost is lower, but that excludes web search cost, so the cost readout should be treated as incomplete.
Matched pairs
98
Companies
61
CLI success
45%
MCP success
33%
Comparison table
| Question | CLI | MCP | Winner | |
|---|---|---|---|---|
Reliability Pass/fail completion rate for the benchmark task. | 45% | 33% | CLI | 21 CLI / 9 MCP / 68 tie |
Throughput Which surface gets agents to an answer faster. | 4m 53s | 7m 2s | CLI | 39 CLI / 18 MCP / 0 tie |
Reported run cost Which surface reports lower run cost before web-search cost. | $0.82 | $0.70 | MCP | 22 CLI / 35 MCP / 0 tie |
Tool load Which surface requires fewer normalized tool calls. | 15 | 15.5 | CLI | 34 CLI / 53 MCP / 11 tie |
Error load Which surface produces fewer errors per run. | 2.3 | 2.9 | CLI | 34 CLI / 38 MCP / 26 tie |
Search overhead Which surface requires less discovery work. | 0.3 | 2.7 | CLI | +2.4 MCP minus CLI |
Investment priorities
What to invest in next
The data points to a practical agent-readiness stack: make official docs easy to find, publish agent-readable context, keep structured API references current, and ship examples that agents can inspect before they start writing code.
Agent-readable infrastructure
The best assets help agents make the first correct request.
These comparisons split API runs by whether a signal was detected or used. Score, runtime, and tool-call deltas show whether the asset improves agent API usage in the current benchmark set.
llms.txt
257 runs with signal · 17 without
With
58
Without
35
Runtime
-36s
Calls
+9
OpenAPI
170 runs with signal · 104 without
With
61
Without
50
Runtime
-11s
Calls
-1
Typed SDK
229 runs with signal · 45 without
With
59
Without
47
Runtime
+18s
Calls
+5
Docs used during run
180 runs with signal · 94 without
With
68
Without
35
Runtime
+1m 53s
Calls
+20
Usage fixes
The fastest improvements are the ones an agent can inspect.
These are not generic DX recommendations. They come from places where agents searched, used docs, hit auth/setup work, or needed recovery.
Make official docs impossible to miss
Runs with Docs used in trace are associated with 33.2 more score points in the current benchmark set.
Add or improve llms.txt
Runs with llms.txt are associated with 23.4 more score points in the current benchmark set.
Add or improve MCP server
Runs with MCP server are associated with 15.2 more score points in the current benchmark set.
Add or improve Agent skills
Runs with Agent skills are associated with 12.5 more score points in the current benchmark set.
Add or improve Typed SDK
Runs with Typed SDK are associated with 12.2 more score points in the current benchmark set.
Add or improve OpenAPI
Runs with OpenAPI are associated with 11.2 more score points in the current benchmark set.
Add or improve Context7
Runs with Context7 are associated with 6.5 more score points in the current benchmark set.
Agent behavior
Claude Code vs Codex
Codex is 27s faster on average in the current API results. The comparison below shows how each agent spends API work across runtime, cost, search, docs, error outputs, and common tool phases.
Runtime shape
The agents spend their work differently.
Search, documentation, execution, and auth/setup calls show where agent users pay the operating cost of an API integration.
Runtime gap
27s
Docs touched
66%
API success rate
75%
Claude Code
137 API runs
Avg score
58Runtime
3m 49s
Cost
$0.38
Calls
17
Search
1
Docs
6
Errors
4
Codex
137 API runs
Avg score
56Runtime
3m 22s
Cost
$1.68
Calls
16
Search
3
Docs
4
Errors
3
API Arena
API
Overall API performance first, then every API category as its own chart.
Overall
Categories
API category charts
Sorted by score
Inference
Payment
Sandboxes
Voice Telephony
Search
Durable Workflow
Cloud Hosting
Stablecoin
Audio
Auth
Vector Databases
Voice STT
Voice TTS
Document Parsing
Meeting Bot
Voice Infra
Agent Memory
Browser
Graph Databases
Agent Automation
E-Signature
Public Sector Intelligence
Unified API
Video Infra
Webscraping
Observability
Storage
Verification
Video Agent
Agent UI
API Testing
Geocoding
Media Generation
Neocloud
Proxy
Sales Intelligence
CLI Arena
CLI
The same benchmark lens applied to command-line integrations and setup flows.
Overall
Categories
CLI category charts
Sorted by score
Payment
Inference
Durable Workflow
Sandboxes
Auth
Search
Vector Databases
Cloud Hosting
Browser
Code Review
Document Parsing
Voice Infra
Voice STT
Voice Telephony
Voice TTS
Neocloud
Observability
Storage
Unified API
MCP Arena
MCP
MCP server results grouped by category so implementation gaps are easier to scan.
Overall
Categories
MCP category charts
Sorted by score
Search
Payment
Voice Telephony
Auth
Vector Databases
Cloud Hosting
Inference
Sandboxes
Agent MCP Gateway
Meeting Bot
Stablecoin
Voice TTS
Document Parsing
Durable Workflow
Voice Infra
Webscraping
Database
Observability
Proxy
Storage
Unified API
Voice STT
AI visibility tools guide
Visibility is only half the agent experience
Compare Lightsage, Profound, Otterly, and AthenaHQ across answer engines, coding agents, and verified API usability.
Full benchmark
View the rest of the results
Continue into DevTool Arena for ranked provider tables, category breakdowns, and company-level evaluation traces.
Join the mailing list
Get updates on leaderboard changes, new benchmark releases, and product announcements.