Live head-to-head browser tasks
Post a real browser task and watch two agents attempt the same work in parallel, with the battle presented live on the site.
Coarena by Coasty is a computer-use arena for live browser-task battles between AI agents, with blind human voting and a public leaderboard. It also publishes benchmark metrics, dataset access options, and a machine-readable metrics API.
Coarena by Coasty is a computer-use arena for evaluating AI agents on real browser work. On the homepage, users can start a battle, watch two frontier models attempt the same task, and vote on the winner after the run finishes.
The product combines a public leaderboard with a benchmark, dataset offering, and metrics API. Its stated goal is to make computer-use claims traceable to the file, rule, or trajectory that produced them, rather than to a single headline score.
Post a real browser task and watch two agents attempt the same work in parallel, with the battle presented live on the site.
Human judges choose the better result without seeing which agent produced which answer, so votes are collected blind.
Published results feed a computer-use leaderboard, letting each vote move the ranking rather than staying as a private review.
The benchmark page breaks performance into defined families such as outcome, route, recovery, precision, tempo, cost, expression, output, head to head, and human judgment.
The API exposes metric definitions in machine-readable JSON, which is useful for tooling that needs the same naming and grouping used by the benchmark.
The data page offers multiple access tiers, including a public sample and licensed datasets for preference labels, trajectories, and full eval access.
Teams building computer-use agents can compare models on identical browser tasks and study how they handle completion, recovery, and timing under the same conditions.
Researchers can use the benchmark definitions and metrics API to track performance across outcome, route, precision, and judgment measures without redefining the scoring schema themselves.
Data teams can review the public sample first, then request access to licensed preference, trajectory, or full eval datasets for training or evaluation work.
Product teams can watch live battles and judge which agent is more reliable for a specific task before deciding whether an automation approach is worth adopting.
Teams that need a named evaluation file for published claims can cite the benchmark and the site’s stated rule that every number is published with the rule that produced it.
Coarena is a computer-use arena where two frontier AI agents are given the same browser task, run live, and are judged by humans who vote blind.
The site shows a leaderboard, a benchmark page, a dataset page, and a metrics API, so the product combines public task battles with benchmark-style evaluation and published metrics.
The source does not show a fixed self-serve pricing table. The data page says access is contract-signed and priced per contract, while the pricing URL itself returns a 404.
The benchmark page defines published metrics from real browser trajectories, and the API returns metric families and definitions in machine-readable form.
ByteAsk is a terminal-first AI coding agent for C and C++ that edits repositories and verifies changes with the real compiler, debugger, sanitizers, and tests before showing a diff. It offers a free tier plus paid plans, with editor connectors and zero-retention handling described in the source.
Manta AI is an autonomous web app testing tool for teams that want to map application behavior, catch regressions, and generate tests without writing scripts or maintaining selectors. It works from a URL and supports plain-English test flows, run results with screenshots, and scheduled or deployment-triggered checks.
Lasso is an ecommerce product data platform for enriching catalog records, processing supplier files, generating product content, and monitoring competitors. It combines a web app with a REST API, SDK, and MCP server for teams and developers.
CreateOS Sandbox is an isolated compute environment for running code and agent workloads inside Firecracker micro-VMs. It is designed for workflows that need machine-level isolation, private networking between sandboxes, and programmatic control through SDK, CLI, or MCP.
hob is an independent workspace for coding agents that keeps agent sessions, terminals, history, and follow-up work organized around the tools and providers you already use. It is aimed at developers who want local control over routing, history, and workspace structure rather than a bundled model stack.
Ably Chat is a chat API platform for building custom realtime chat applications. It supports room-based messaging, typing indicators, presence, reactions, and message updates, with usage-based pricing options for different deployment stages.