Testbed is a robotics research and product company. We are building a future where anyone can build, train, and improve their own physical AI, whatever their expertise.

Robotics has a small number of people who know how to make a machine better, and almost everyone else. That knowledge sits inside a few companies and a few labs, and it moves by apprenticeship rather than by documentation. A team that wants to improve a robot needs a statistician, a fleet, and a year. Most teams have none of the three.

Grace Hopper wrote the first compiler and could not give it away. “I had a running compiler and nobody would touch it… they carefully told me, computers could only do arithmetic; they could not do programs.” The compiler was not a faster machine. It was a translation layer that let people who were not machine-code specialists write software, and programming stopped being a priesthood. Robotics has not had that moment.

We start with the decision, not the machine.The hardest expertise to hire in robotics is not building a robot. It is knowing whether the version built this month is better than the one shipped last month — on real hardware, under real conditions, where the numbers are noisy and the machines are not identical. Teams answer that today with a review meeting and somebody's judgement.

The release gate is our first product. Put a new build on a few machines and leave the rest on the old one. Both groups already report how they are doing. We compare them and return one verdict, with the arithmetic attached, before the build reaches the rest of the fleet. Nothing stops working and nothing goes into a shop.

fleet · 40 machines
canary · 5 · v2.4.0baseline · 35 · v2.3.1
evidence4h windows

window 0 · 0 device-hours · e-stop baseline
not decisive

not enough evidence yet

0 of ~120 device-hours. Leaning promote.
You'd know by Thu 14:00 — or Tue 09:00 with 6 more machines.

nothing is promoted until it clears
example · synthetic data

Measurement should not require a statistician. The mathematics of comparing two noisy populations is well understood, and almost nobody outside a large lab has it on staff. We are packaging it so a team of four gets the answer a team of four hundred would get.

Everything replays. The same inputs and the same pinned version always produce the same verdict, byte for byte. A decision made in March can be re-derived in September from the raw data. Anything that gates a physical machine should be auditable by the people it gates.

What we learn, we publish.The thing nobody can look up is how jumpy a robot's numbers are normally — how much they swing between sites, between shifts, and as machines wear. That is what decides how long anyone has to wait for an answer, and it is not a trade secret. We intend to make it common knowledge rather than an advantage.

What we have not earned the right to say.We have not run on a customer's fleet. Every rate we can quote today rests on synthetic data. We cannot claim a team runs fewer physical trials, because that would require showing our comparison agrees with what a real trial would have concluded, and we have never measured it. Until that changes, the honest version is narrower: find out from the fleet you already run.

We are looking for people who ship software to real machines and have to decide whether the new version is better. If that is the job, we would like to hear how it is done today.