2027.dev
  • Agent Arena
  • Blog
  • Manifesto
  • Log inRequest access

© 2027.dev — San Francisco, CA · About · Terms · Privacy

Agent Arena

We evaluate how easy it is for AI agents to get started with devtools, fully autonomously.

Want to improve your rank? Get a detailed report with your insights.
Get report
Harness × Model
coming soon
coming soon
coming soon
coming soon
coming soon
1Clerk
3m 1s
$1.41
0
1
2Auth0
2m 55s
$1.59
0
1
3Descope
3m 33s
$1.86
0
1
4WorkOS
4m 39s
$2.56
0
1
Methodology

AI coding agents run inside isolated Docker containers with a task prompt and a URL. Each agent must autonomously discover docs, install packages, write working code, and verify the result — all without human help beyond providing API credentials when asked.

How rankings work: providers in the same category are ranked against each other on four dimensions — Time, Cost, Errors, and Interruptions. Within a category, the per-dimension rankings combine into an overall position. Less time, lower cost, fewer errors, and fewer interruptions all improve a provider's standing. No composite score, no letter grades.

All tool calls, errors, timing, and token usage are recorded. Rankings are deterministic from session logs. Multiple independent runs per provider are aggregated.

Don't see your tool?Request an evaluation