HomeCompaniesCoasty

Real World Evals and RL Environments for Computer-Use

Coasty is a computer use agent that officially reached 82.8% on OSWorld, which puts us among the most reliable agents out there. Going through the evaluation process, we realized a better crowdsourced benchmark was needed. So we created Coarena, real-world evals and RL environments for computer-use agents. We run a live arena where people post the computer-use tasks they actually need done, two frontier models race them head-to-head, and humans judge the outcome blind. Every battle produces something labs can't get from a static benchmark, an eval built on real demand instead of a memorized test set. Most benchmarks demo well and break the moment a model trains on them. They're fixed, they grade only the final screen, and they miss everything that matters, the run that spent $240 on an "under $200" task, the agent that subscribed to three newsletters on the way to checkout, the one that looped for 40 steps and quit. OSWorld and WebArena saturate in months. Ours are parameterized and demand-sourced, so the scores stay honest. On top of the arena we build deterministic RL environments - flights, commerce, CRM, forms that reset bit-for-bit and grade outcome, trajectory, side-effects, and constraints, with per-step failure attribution and a dense reward you can train on. We also grade what nobody else does: whether an agent can be hijacked by instructions hidden in a page, and whether it knows when to stop and ask instead of spending someone's money.
Active Founders
Prateek Jannu
Prateek Jannu
Founder
Building state-of-the-art computer-use agents @ coasty.ai. Enterprise-grade automation for work that still runs on humans clicking buttons. Stanford dropout, Purdue before.
Nitish Kovuru
Nitish Kovuru
Founder
Co-founder of Coasty, SOTA CUA framework on OSWorld at 82% accuracy. Did CS at Columbia (Vision track) and hold a BS in Computer Engineering from Purdue, where I TA'd a graduate-level AI course as an undergrad and later went on to build enterprise sales automation agentic systems for my previous company.
Company Launches
Coasty: The computer-use agent that doesn't break on real software
See original launch post

TL;DR: Coasty gives developers a computer-use agent that reliably automates back-office work in legacy desktop software where RPA scripts and other CUAs fail. It’s achieved 82.81% on OSWorld and it’s great at automating workflows where quiet mistakes cost the most: regulatory filings, patient records, payments, and gives developers precise control and observability over its work.

https://www.youtube.com/watch?v=3EwLfqlgDEM\

(The portals and the data in it, are mockups)

Replayable audit trail (real run): https://coasty.ai/share/5b3320fc-0e64-4f1b-b684-9fd6e8cd8960 

OSWorld submission (Coasty scored 82.81% on OSWorld Verified, confirmed by the OSWorld team, with our latest internal run at 85.60%.): https://github.com/coasty-ai/coasty-osworld

The Problem

Most back-office work doesn't happen in clean web apps. It happens in legacy desktop software, full of surprise pop-ups, wandering menus, and impenetrable UIs.

Today, developers automating this back-office “real” work are stuck choosing between brittle RPA scripts and CUAs that demo beautifully and score high but break on contact.

The problem isn't getting an agent to click. It's trusting it when a silent mistake gets someone fired.

Why Coasty is reliable where others aren't

It recovers mid-task. Coasty catches its errors and fixes them mid-run. If it’s crunching on a 20-step workflow and makes a misclick at step 3, it will recover and start from the error. First-pass accuracy is vanity. Recovery keeps a run alive.

It survives messy software. Pop-ups, modals that steal focus, legacy desktop apps. Coasty doesn’t record clicks, it reads the screen. So when the UI changes overnight and every RPA script breaks, Coasty looks at the new screen and keeps going.

It knows when it's done. An agent that thinks it finished but didn't is worse than one that fails loudly. Coasty verifies its own output and leaves a timestamped audit trail you can replay.

It swarms. One task across forty state portals, each with its own ancient interface, all at once. Forty agents, one verified result set. What takes a team a week, is done overnight.

Traction

Coasty has been verified by the OSWorld team and will be published soon, but also another internal benchmark we published (85.60%): <https://github.com/coasty-ai/coasty-osworld >.

We have paying customers on four continents, a few thousand active users, and we have actually grown 24x since Claude Cowork came out.

Why us

We started working on this idea when models weren’t good enough for computer use (last summer), so we’ve built frameworks to test and run the perfect combination of models and harnesses to see what works in practice. Prateek worked as an ML engineer at multiple startups and dropped his Stanford admit while I did CS at Columbia and built my previous company’s entire agentic sales automation. We’ve grown with this field and have been working on it even before it got much attention, so we understand where CUAs mess up and how to deal with these messes realistically.

The ask

Send us one workflow your team runs by hand. We’ll tell you straight whether Coasty can run it reliably today. If it looks automatable, we’ll build a working proof of concept on your real software before a long sales process.

We’re especially looking for intros to operators and executives in accounting, tax, insurance, healthcare, compliance, and data privacy - anywhere teams are still babysitting RPAs, clicking through legacy portals, reconciling records, filing forms, or moving data across old desktop software.

Here's our Calendly: https://cal.com/coasty/15min

Coasty
Founded:2026
Batch:Summer 2026
Team Size:2
Status:
Active
Location:San Francisco
Primary Partner:Brad Flora