# Agent Experience Arena

The Agent Experience Arena combines agent discovery and agent usability into one developer-tool ranking.

Default view: **Claude Code x all contexts**.

Discovery snapshot: 2026-06-13.

## Scoring

- **Discovery score** uses the current visibility score from [/leaderboard/discoverability](/leaderboard/discoverability).
- **Usability score** averages current API, CLI, and MCP integration scores from the selected agent's DevTool Arena leaderboards.
- **Agent Experience score** is currently an equal blend of discovery and usability.
- Included companies must have both discoverability evidence and at least one scored API, CLI, or MCP usability eval for the selected agent.

## Combined Ranking

| Rank | Tool | Category | Experience | Discovery | Usability | Visibility | Avg Rank | Best Rank | API | CLI | MCP |
|------|------|----------|------------|-----------|-----------|------------|----------|-----------|-----|-----|-----|
| 1 | ElevenLabs | Voice TTS | 87 | 94 | 80 | 94% | 1.3 | 1 | 81 | 68 | 92 |
| 2 | Stripe | Payment | 84 | 87 | 80 | 87% | 1.5 | 1 | 81 | 71 | 87 |
| 3 | Browserbase | Browser | 83 | 94 | 71 | 94% | 4.0 | 4 | 72 | 69 | — |
| 4 | Browser Use | Browser | 78 | 90 | 66 | 90% | 6.0 | 6 | 62 | 69 | — |
| 5 | Pinecone | Vector Databases | 73 | 74 | 71 | 74% | 3.3 | 3 | 65 | 68 | 80 |
| 6 | Twilio | Voice Telephony | 73 | 92 | 53 | 92% | 1.1 | 1 | 62 | 65 | 33 |
| 7 | Resend | Email | 72 | 77 | 66 | 77% | 1.4 | 1 | 65 | 64 | 69 |
| 8 | Circle | Stablecoin | 68 | 74 | 62 | 74% | 3.1 | 1 | 62 | — | — |
| 9 | Temporal | Durable Workflow | 67 | 54 | 79 | 54% | 2.2 | 1 | 86 | 72 | — |
| 10 | Clerk | Auth | 66 | 58 | 74 | 58% | 2.0 | 1 | 67 | — | 81 |
| 11 | LiveKit | Voice Infra | 65 | 70 | 60 | 70% | 1.7 | 1 | 86 | 88 | 6 |
| 12 | Vonage | Voice Telephony | 65 | 68 | 62 | 68% | 2.9 | 2 | — | 62 | — |
| 13 | Merge | Unified API | 64 | 58 | 69 | 58% | 2.4 | 1 | 69 | — | — |
| 14 | Deepgram | Voice STT | 63 | 76 | 49 | 76% | 3.6 | 1 | 72 | 70 | 6 |
| 15 | Qdrant | Vector Databases | 62 | 73 | 51 | 73% | 2.6 | 1 | 74 | — | 27 |
| 16 | Modal | Sandboxes | 61 | 47 | 74 | 47% | 3.2 | 3 | 81 | 67 | — |
| 17 | Daytona | Sandboxes | 60 | 61 | 59 | 61% | 1.7 | 1 | 77 | 89 | 10 |
| 18 | E2B | Sandboxes | 60 | 63 | 56 | 63% | 1.6 | 1 | 68 | 75 | 25 |
| 19 | Recall.ai | Meeting Bot | 59 | 75 | 43 | 75% | 1.6 | 1 | 70 | — | 15 |
| 20 | Datadog | Observability | 58 | 65 | 50 | 65% | 3.9 | 3 | 76 | 70 | 3 |
| 21 | AssemblyAI | Voice STT | 57 | 76 | 38 | 76% | 3.6 | 1 | 69 | 6 | — |
| 22 | LemonSqueezy | Payment | 56 | 35 | 76 | 35% | 3.2 | 2 | 76 | — | — |
| 23 | Unstructured | Document Parsing | 56 | 46 | 66 | 46% | 3.2 | 2 | 63 | 68 | — |
| 24 | Daily | Voice Infra | 55 | 46 | 63 | 46% | 3.9 | 2 | 63 | — | — |
| 25 | OpenCageData | Geocoding | 55 | 21 | 89 | 21% | 4.6 | 2 | 89 | — | — |
| 26 | Tavily | Search | 55 | 40 | 70 | 40% | 2.2 | 1 | 69 | 69 | 73 |
| 27 | Exa | Search | 53 | 34 | 71 | 34% | 3.5 | 2 | 70 | — | 71 |
| 28 | Render | Cloud Hosting | 53 | 54 | 52 | 54% | 3.2 | 1 | 65 | 9 | 81 |
| 29 | WorkOS | Auth | 52 | 35 | 68 | 35% | 3.1 | 1 | 67 | 68 | — |
| 30 | Apideck | Unified API | 51 | 35 | 66 | 35% | 3.7 | 2 | 66 | — | — |
| 31 | LlamaParse | Document Parsing | 51 | 32 | 70 | 32% | 2.7 | 1 | 70 | 69 | — |
| 32 | PayPal | Payment | 51 | 25 | 76 | 25% | 3.9 | 2 | 69 | — | 83 |
| 33 | Firecrawl | Search | 50 | 31 | 68 | 31% | 3.9 | 1 | 65 | 70 | 70 |
| 34 | Rev AI | Voice STT | 50 | 14 | 86 | 14% | 5.8 | 3 | 86 | — | — |
| 35 | Agora | Voice Infra | 49 | 34 | 63 | 34% | 3.3 | 2 | 63 | — | — |
| 36 | Plivo | Voice Telephony | 49 | 50 | 47 | 50% | 4.2 | 3 | 56 | — | 38 |
| 37 | Chroma | Vector Databases | 48 | 44 | 51 | 44% | 4.7 | 3 | 77 | 3 | 73 |
| 38 | Speechify | Voice TTS | 48 | 0 | 96 | 0% | — | — | 96 | — | — |
| 39 | Vercel | Sandboxes | 48 | 39 | 56 | 39% | 2.5 | 1 | 43 | 68 | — |
| 40 | Jina AI | Search | 47 | 15 | 78 | 15% | 5.2 | 2 | 75 | 77 | 81 |
| 41 | LanceDB | Vector Databases | 46 | 11 | 81 | 11% | 6.3 | 2 | 78 | — | 84 |
| 42 | Auth0 | Auth | 45 | 40 | 49 | 40% | 3.3 | 2 | 62 | 81 | 3 |
| 43 | Sprites | Sandboxes | 45 | 1 | 88 | 1% | 5.5 | 4 | — | 88 | — |
| 44 | Paddle | Payment | 44 | 41 | 46 | 41% | 3.0 | 1 | 64 | — | 28 |
| 45 | Railway | Cloud Hosting | 44 | 54 | 33 | 54% | 3.0 | 2 | 62 | 3 | — |
| 46 | LMNT | Voice TTS | 43 | 0 | 86 | 0% | 6.0 | 6 | 86 | — | — |
| 47 | Meetstream | Meeting Bot | 43 | 4 | 81 | 4% | 2.8 | 2 | 81 | — | — |
| 48 | Mollie | Payment | 43 | 0 | 85 | 0% | — | — | 85 | — | — |
| 49 | Pioneer AI | Inference | 42 | 0 | 83 | 0% | — | — | 83 | — | — |
| 50 | Reducto | Document Parsing | 42 | 12 | 72 | 12% | 3.5 | 2 | 78 | 70 | 68 |
| 51 | Telnyx | Voice Telephony | 42 | 43 | 40 | 43% | 2.5 | 2 | 63 | 27 | 31 |
| 52 | Chargebee | Payment | 41 | 18 | 63 | 18% | 4.0 | 3 | 63 | — | — |
| 53 | Prelude | Verification | 41 | 1 | 81 | 1% | 5.0 | 5 | 81 | — | — |
| 54 | Rime | Voice TTS | 41 | 0 | 81 | 0% | — | — | 96 | 65 | — |
| 55 | Sonilo | Audio | 41 | 0 | 81 | 0% | — | — | 81 | — | — |
| 56 | Box | Storage | 40 | 0 | 79 | 0% | — | — | 85 | 69 | 83 |
| 57 | Merge Agent Handler | Agent MCP Gateway | 40 | 0 | 80 | 0% | — | — | — | — | 80 |
| 58 | You.com | Search | 40 | 5 | 75 | 5% | 5.8 | 4 | 68 | — | 81 |
| 59 | BetterStack | Observability | 39 | 15 | 63 | 15% | 6.3 | 5 | 63 | — | — |
| 60 | Superserve | Sandboxes | 39 | 0 | 78 | 0% | — | — | 78 | — | — |
| 61 | Surge | Verification | 39 | 0 | 78 | 0% | — | — | 78 | — | — |
| 62 | Anchor Browser | Browser | 38 | 0 | 76 | 0% | — | — | 77 | 74 | — |
| 63 | Composio | Agent MCP Gateway | 38 | 0 | 75 | 0% | — | — | — | — | 75 |
| 64 | DFNS | Stablecoin | 38 | 1 | 75 | 1% | 10.0 | 10 | 75 | — | — |
| 65 | MeetGeek | Meeting Bot | 38 | 1 | 75 | 1% | 4.0 | 4 | 62 | — | 88 |
| 66 | Memgraph | Graph Databases | 38 | 0 | 76 | 0% | — | — | 76 | — | — |
| 67 | Neo4j | Graph Databases | 38 | 0 | 76 | 0% | — | — | 76 | — | — |
| 68 | OpenRouter | Inference | 38 | 9 | 67 | 9% | 4.3 | 2 | 70 | 63 | — |
| 69 | Postman | API Testing | 38 | 0 | 75 | 0% | — | — | 75 | — | — |
| 70 | Prefect | Durable Workflow | 38 | 17 | 59 | 17% | 3.6 | 2 | 67 | — | 51 |
| 71 | Replicate | Inference | 38 | 5 | 70 | 5% | 4.2 | 4 | 70 | — | — |
| 72 | Agentmail | Email | 37 | 0 | 73 | 0% | — | — | 67 | 75 | 77 |
| 73 | ElevenLabs Music | Audio | 37 | 0 | 74 | 0% | — | — | 74 | — | — |
| 74 | Mem0 | Vector Databases | 37 | 13 | 61 | 13% | 2.1 | 2 | 61 | — | — |
| 75 | Stytch | Auth | 37 | 11 | 63 | 11% | 4.7 | 2 | 63 | — | — |
| 76 | Descope | Auth | 36 | 0 | 72 | 0% | 4.0 | 4 | 68 | 65 | 84 |
| 77 | Docusign | E-Signature | 36 | 0 | 71 | 0% | — | — | 71 | — | — |
| 78 | Restate | Durable Workflow | 36 | 4 | 68 | 4% | 2.6 | 2 | 68 | — | — |
| 79 | Vapi | Voice Telephony | 36 | 11 | 60 | 11% | 3.3 | 3 | 60 | — | — |
| 80 | Zep | Vector Databases | 36 | 9 | 62 | 9% | 3.5 | 3 | 62 | — | — |
| 81 | AIsa | Inference | 35 | 0 | 69 | 0% | — | — | 69 | — | — |
| 82 | Dropbox | Storage | 35 | 0 | 70 | 0% | — | — | 70 | — | — |
| 83 | Google DeepMind | Inference | 35 | 0 | 70 | 0% | — | — | 70 | — | — |
| 84 | Nebius Token Factory | Inference | 35 | 0 | 70 | 0% | — | — | 70 | — | — |
| 85 | PandaDoc | E-Signature | 35 | 0 | 70 | 0% | — | — | 70 | — | — |
| 86 | SambaNova | Inference | 35 | 1 | 69 | 1% | 3.0 | 3 | 69 | — | — |
| 87 | Steel | Browser | 35 | 0 | 70 | 0% | — | — | 65 | 74 | — |
| 88 | Tenki | Sandboxes | 35 | 0 | 70 | 0% | — | — | 65 | 74 | — |
| 89 | Cartesia | Voice TTS | 34 | 18 | 49 | 18% | 3.6 | 2 | 70 | 6 | 71 |
| 90 | Diffbot | Webscraping | 34 | 2 | 65 | 2% | 4.8 | 4 | 65 | — | — |
| 91 | Extend.ai | Document Parsing | 34 | 0 | 68 | 0% | — | — | 67 | 67 | 69 |
| 92 | Google Lyria | Audio | 34 | 0 | 67 | 0% | — | — | 67 | — | — |
| 93 | HitPay | Payment | 34 | 0 | 68 | 0% | — | — | 68 | — | — |
| 94 | Infobip | Voice Telephony | 34 | 3 | 65 | 3% | 2.5 | 2 | 65 | — | — |
| 95 | Meeting BaaS | Meeting Bot | 34 | 34 | 34 | 34% | 3.1 | 2 | 66 | — | 1 |
| 96 | Tavus | Video Agent | 34 | 0 | 68 | 0% | — | — | 68 | — | — |
| 97 | TinyFish | Search | 34 | 0 | 68 | 0% | 4.0 | 4 | 68 | — | — |
| 98 | BoldSign | E-Signature | 33 | 0 | 65 | 0% | — | — | 65 | — | — |
| 99 | FalkorDB | Graph Databases | 33 | 0 | 65 | 0% | — | — | 65 | — | — |
| 100 | Loudly | Audio | 33 | 0 | 65 | 0% | — | — | 65 | — | — |
| 101 | Merge Gateway | Inference | 33 | 0 | 65 | 0% | — | — | 65 | — | — |
| 102 | Relevance AI | Agent Automation | 33 | 0 | 65 | 0% | — | — | 65 | — | — |
| 103 | TigerGraph | Graph Databases | 33 | 0 | 65 | 0% | — | — | 65 | — | — |
| 104 | Allo | Voice Telephony | 32 | 0 | 64 | 0% | — | — | 64 | — | — |
| 105 | Blaxel | Sandboxes | 32 | 2 | 62 | 2% | 5.3 | 4 | 62 | — | — |
| 106 | Cloudflare | Cloud Hosting | 32 | 19 | 44 | 19% | 4.0 | 1 | 63 | 65 | 4 |
| 107 | Kite | Stablecoin | 32 | 0 | 63 | 0% | — | — | 63 | — | — |
| 108 | Nimble | Search | 32 | 0 | 64 | 0% | 6.0 | 6 | 62 | 63 | 68 |
| 109 | Sensible | Document Parsing | 32 | 1 | 62 | 1% | 9.0 | 9 | 62 | — | — |
| 110 | Cognee | Agent Memory | 31 | 0 | 61 | 0% | — | — | 61 | — | — |
| 111 | Dust | Agent Automation | 31 | 0 | 62 | 0% | — | — | 62 | — | — |
| 112 | Gumloop | Agent Automation | 31 | 0 | 62 | 0% | — | — | 62 | — | — |
| 113 | HeyGen | Video Agent | 31 | 0 | 62 | 0% | — | — | 62 | — | — |
| 114 | Letta | Agent Memory | 31 | 0 | 61 | 0% | — | — | 61 | — | — |
| 115 | Netlify | Cloud Hosting | 31 | 44 | 17 | 44% | 3.3 | 2 | — | 4 | 29 |
| 116 | Pursuit | Public Sector Intelligence | 31 | 0 | 62 | 0% | — | — | 62 | — | — |
| 117 | Weaviate | Vector Databases | 31 | 60 | 2 | 60% | 4.0 | 3 | — | — | 2 |
| 118 | Zoom AI Services | Voice STT | 31 | 0 | 61 | 0% | — | — | 61 | — | — |
| 119 | Resemble AI | Voice TTS | 30 | 1 | 58 | 1% | 5.5 | 5 | 83 | — | 33 |
| 120 | Fireworks AI | Inference | 29 | 11 | 46 | 11% | 2.4 | 2 | 84 | 7 | — |
| 121 | Convex | Database | 28 | 51 | 5 | 51% | 2.0 | 1 | — | — | 5 |
| 122 | Zilliz Cloud | Vector Databases | 27 | 9 | 45 | 9% | 5.2 | 2 | 62 | — | 28 |
| 123 | Groq | Inference | 26 | 18 | 33 | 18% | 3.5 | 1 | 90 | 4 | 5 |
| 124 | Cerebras | Inference | 25 | 3 | 47 | 3% | 4.1 | 1 | 78 | — | 16 |
| 125 | fastCRW | Webscraping | 25 | 0 | 50 | 0% | — | — | 65 | — | 35 |
| 126 | Nebius AI Cloud | Neocloud | 25 | 0 | 49 | 0% | — | — | 66 | 31 | — |
| 127 | Scalekit | Auth | 24 | 2 | 45 | 2% | 4.8 | 3 | 62 | 32 | 42 |
| 128 | Scrapfly | Webscraping | 22 | 0 | 43 | 0% | — | — | 81 | — | 5 |
| 129 | CodeRabbit | Code Review | 20 | 33 | 7 | 33% | 1.8 | 1 | — | 7 | — |
| 130 | Massive | Proxy | 20 | 5 | 34 | 5% | 9.2 | 7 | 63 | — | 5 |
| 131 | Naive | Unified API | 20 | 0 | 40 | 0% | — | — | 76 | 38 | 7 |
| 132 | Razorpay | Payment | 20 | 0 | 39 | 0% | 4.0 | 4 | 74 | — | 3 |
| 133 | Mubert | Audio | 18 | 0 | 35 | 0% | — | — | 35 | — | — |
| 134 | Square | Payment | 16 | 3 | 28 | 3% | 5.4 | 4 | — | — | 28 |
| 135 | NationGraph | Public Sector Intelligence | 15 | 0 | 29 | 0% | — | — | 29 | — | — |
| 136 | BlindPay | Stablecoin | 14 | 0 | 28 | 0% | — | — | 27 | — | 29 |
| 137 | Freestyle | Sandboxes | 14 | 1 | 27 | 1% | 4.0 | 4 | 27 | — | — |
| 138 | GovSpend | Public Sector Intelligence | 14 | 0 | 28 | 0% | — | — | 28 | — | — |
| 139 | CopilotKit | Agent UI | 13 | 0 | 26 | 0% | — | — | 26 | — | — |
| 140 | Semgrep | Code Review | 13 | 16 | 9 | 16% | 3.5 | 3 | — | 9 | — |
| 141 | Arcade | Agent MCP Gateway | 12 | 0 | 23 | 0% | — | — | — | — | 23 |
| 142 | Fireblocks | Stablecoin | 12 | 21 | 2 | 21% | 4.9 | 2 | — | — | 2 |
| 143 | Stability AI | Audio | 12 | 0 | 23 | 0% | — | — | 23 | — | — |
| 144 | Coinbase Payments | Stablecoin | 11 | 7 | 14 | 7% | 3.9 | 2 | 26 | — | 1 |
| 145 | Upsun | Cloud Hosting | 11 | 1 | 21 | 1% | 6.0 | 6 | 21 | — | — |
| 146 | TABStack | Search | 8 | 0 | 16 | 0% | — | — | 27 | — | 5 |
| 147 | Greptile | Code Review | 7 | 7 | 7 | 7% | 4.1 | 2 | — | 7 | — |
| 148 | Qodo | Code Review | 7 | 11 | 3 | 11% | 3.0 | 1 | — | 3 | — |
| 149 | Nylas | Email | 6 | 0 | 12 | 0% | — | — | 21 | — | 3 |
| 150 | Camunda | Durable Workflow | 3 | 3 | 2 | 3% | 3.5 | 3 | — | 2 | — |

## Detail Model

The rendered page includes a market-position landscape, sortable combined ranking table, and row detail modal. The modal separates discovery evidence from usability evidence so agents and humans can inspect why a tool ranked where it did.

## Frequently Asked Questions

### What is the Agent Experience Arena?

The Agent Experience Arena ranks developer tools by how well they work with AI coding agents. It combines two signals into a single score: how often agents discover a tool when solving tasks in its category, and how reliably agents can actually use the tool once they find it.

Unlike a single-task benchmark, the arena aggregates results across many prompts, coding agents, and project contexts so the ranking reflects real-world agent behavior rather than a one-off run.

### What does the discovery score measure?

The discovery score measures how often coding agents surface a tool when asked to solve tasks in its category, across many prompts, agents, and contexts. We track whether the tool appears at all (visibility) and how highly it ranks among the agent's recommendations (average and best rank).

A higher discovery score means agents reach for the tool more consistently and rank it near the top when they do.

### What does the usability score measure?

The usability score measures how reliably a coding agent can integrate and execute real workflows using a tool's API, CLI, and MCP server. We give the agent a practical integration task on each surface and evaluate whether it can complete the job end to end — finding the right docs, generating valid code, and running it successfully.

The usability score blends the tool's API, CLI, and MCP results, so tools that are easy to use across every surface score highest.

### How is the Agent Experience Score calculated?

The Agent Experience Score blends the discovery score and the usability score equally — 50% each. Being easy to find matters just as much as being easy to use, so a tool needs to do well on both to top the arena.

Scores map to letter grades: A (85+), B (72–84), C (58–71), and D below 58.

### How do the coding agent and context filters work?

You can switch the active coding agent to see how discovery and usability change from one agent to another — the same tool can rank differently depending on which agent is asking. Scores are aggregated across the project contexts we test (such as Next.js and FastAPI) so the ranking reflects a range of real integration scenarios rather than a single stack.

Canonical URL: https://lightsage.com/agent-experience-arena
