We’re Jay and Aris, co-founders of Robocurve.
Robots can do karate kicks, cut carrots, and bake bread in demo videos. But how capable are they, really? Robocurve measures how well AI and robots can do real-world tasks. We are a Public Benefit Corporation and our work is open-source.
Ask:
🤖 Visit robocurve.org to get your AI model or robot evaluated
⭐ Inspect Robots is open source (inspectrobots.org). You can use it to evaluate your own models and tell us how it goes.
Frontier labs are targeting general-purpose robotics by 2028, yet the field of robotics evals barely exists. Labs evaluate in-house, and no independent group has stepped in to run continuous benchmarking as a service. The result is that no one actually knows how good anyone else is, or where the real frontier sits.
Good benchmarks require operating and maintaining physical hardware and real-world setups. Simulation only goes so far, since models can show large performance gaps in the real world. Building real-world benchmarks means coordinating job-domain experts, evals engineering, and hands-on robotics all at once.
Part 1: Open-source evaluation tooling
We are scaling the field of robotics evals as fast as possible by building open-source tools and benchmarks. We are releasing Inspect Robots, an open-source framework to evaluate VLAs and LLMs on robots, with full transcript logging, tracebacks, and Rerun support.
Using Inspect Robots, we found that Claude Fable 5 could control YAM arms to complete simple tasks, sometimes outperforming state-of-the-art open-source VLAs like MolmoAct2. We will post the full results soon.
Inspect Robots is real-world first with digital twin simulation support.
Inspect Robots works with 40 VLAs (π0.5, GR00T N1.7, MolmoAct2 and more via XPolicyLab) and 260+ frontier LLMs on day one.
Inspect Robots is free and fully open-source, released under the MIT license.
Part 2: Real-world benchmarks across robots and industries
We are using Inspect Robots to build out multiple high-quality, decision-relevant benchmarks for AI and robots across industries, focusing on those with the highest impact on jobs and AI timelines, so that society can prepare. We welcome everyone from the community to use Inspect Robots for free to build their own benchmarks.
Whoever builds the trusted measure of robot capability becomes the reference everyone relies on. METR, Epoch, and Artificial Analysis have done this for LLMs. We are building Robocurve for physical AI.
Backstory
We combine backgrounds in AI evals and robotics to grow the field of robotics evals as fast as possible.
Jay comes from a background in AI safety (MATS and UK AI Security Institute) and AI engineering (AWS Bedrock). He was motivated by research from Jan Kulveit, Raymond Douglas, Rudolf Laine, Luke Drago, and Jacob Steinhardt on measuring the dynamics between AI and society. Robocurve began as an idea in one of his papers accepted into ACL. Jay was a top contributor to UK AISI’s Inspect Evals, which directly informed the engineering of Inspect Robots.
Aris has a robotics background from Harvard’s Computational Robotics Lab, with research published in IEEE Robotics and Automation Letters. She has engineering experience from Amazon Robotics, Amazon AGI Lab and Yondu Robotics (YC W24), and loves working with robots.
Ask again:
🤖 Visit robocurve.org to get your AI model or robot evaluated
⭐ Inspect Robots is open source (inspectrobots.org). You can use it to evaluate your own models and tell us how it goes.