Ship fleet updates without betting the whole fleet.

Ship fleet updates without betting the whole fleet.

Testbed checks whether a new build is safe before it reaches every robot.

Run it on your own rollout telemetry, at zero risk.

Testbed checks whether a new build is safe before it reaches every robot. Run it on your own rollout telemetry, at zero risk.

See how it works
Gate receipt rcpt_0412 ยท shadow mode
v4.17.3 Baseline
v4.18.0-rc2 Candidate
Canary
12 of 212 robots, 30 min
Policy
fleet-safety.yaml at a3f9c2e
Comparison grade
Regression

Candidate is worse on e-stop rate (0.19, limit 0.05).

Pinned signalBaselineCandidateLimitResult
E-stop rate, per hour0.020.19 +0.17at most 0.05 over
Localization drift0.4 cm0.5 cm +0.1 cmat most 1.0 cmok
Task success99.1%98.9% -0.2%at least 98.5%ok
Control loop p9541 ms43 ms +2 msat most 60 msok
Verdict Pause

Flagged at 14:07. E-stop rate over limit for 9 minutes. Rollout held at 12 of 212.

Example receipt. Same inputs, same verdict. Replayable.

A single bad build can take down an entire fleet.

One flawed update reaches every robot in minutes. Here's what that's already cost the most advanced fleets in the world.

A white Cruise robotaxi with a roof sensor array on a San Francisco street

Autonomous vehicles

Cruise

950

robotaxis recalled

Bad build. Whole fleet. Fixed after the harm.

After hitting a pedestrian, the software tried to pull over and dragged her. All 950 robotaxis were recalled to fix it.

TechCrunch, Nov 2023

Photo: Dllu, CC BY-SA 4.0

Gate

Measure the candidate against the incumbent.

Pin the signals. Canary the new build. The judge compares both arms and writes a receipt.

Scores and signals
v4.17.3 Incumbent v4.18.0-rc2 Candidate
Comparison grade
Regression
Candidate is worse on e-stop rate, bumps per device-hour, and task success.
E-stop rate
0.02 avg
0.19 +0.17
9
Bumps / device-hour
1.1 avg
6.25 +5.15
4
Task success
99.1% avg
98.9% -0.2%
2
Localization drift
0.4 cm avg
0.5 cm +0.1 cm
1
Control loop p95
41 ms avg
43 ms +2 ms
1

The release lifecycle

From "this build should be fine" to "this build is safe to ship"

  1. 01

    Capture a rollout

    Replay a past or live deployment from your fleet's own telemetry.

    Show our engine a rollout
  2. 02

    Define the checks

    Pin the signals that matter, e-stops, interventions, task success, before the test runs.

    Set your checks
  3. 03

    Compare candidate vs. baseline

    Run the new build against the version it's replacing. Side-by-side, over the same window of real fleet data. Nothing is learned, so nothing goes stale.

    Run a Backtest
  4. 04

    Gate the deployment

    Promote only when the evidence clears. A bad build stops at the canary, not the fleet.

    Set up a release gate

The judge scores the pinned signals, every window.

95% Task success
62% E-stop windows
Localization

Automated scorers

Deterministic checks for format, thresholds, and the pinned policy. No model on the verdict path.

FAQ

What is a Backtest?

Show our engine a past rollout from your fleet. It replays the deployment and shows where it would have held a bad build, on your own data, at zero risk. You get a receipt: candidate vs. the previous version, the timestamp of the flag, and how long the signal sat over its limit before the hold. Detection takes a window of telemetry, not an instant.

Do I have to send you our data?

No. You show our engine a rollout. You're not handing over your systems. The critical decision can run on your side, the cloud is optional, and we see your metadata, not your data.

What telemetry do you need?

Only the safety-critical signals you pin in the policy: things like e-stops, interventions, task success, and control-loop metrics. We do not ingest or retain your full telemetry stream. The gate reads the pinned signals over the canary window and keeps the receipt.

What do I have to set up?

There is no custom integration, but there is wiring. You run our adapter against your telemetry, map your signal names to the ones the policy uses, set up authentication, and hook the gate into your updater so it can promote, pause, or roll back. After that it runs on the OTA path you already have.

What is the baseline? Does it go stale?

The baseline is the build you're replacing, measured over the same window as the candidate. Testbed does not learn a model of normal and keep it around, so there is no baseline to retune between releases. Every gate re-measures both builds, and the limits are the ones you pinned in the policy file.

Does an AI model decide whether my build ships?

No. No model sits on the decision path. Every verdict comes from rules you pin before the test, so it's deterministic and auditable.

How is this different from a dashboard?

A dashboard shows you what happened after you deploy. Testbed decides before the build reaches the fleet, and can block a bad one. It's a gate, not a report.

What does "Shadow, Guard, Unguarded" mean?

Autonomy grows in stages. Shadow: Testbed watches and reports. Guard: it can hold a bad rollout. Unguarded: it runs the gate automatically. You move up as you trust it.

How do I get started?

Join the waitlist and we'll run a Backtest on a past rollout from your fleet.

What does it cost?

We're working with early teams now. Reach out and we'll figure out what makes sense.

Who is this for?

Teams shipping software to fleets of physical robots, where a bad deploy crashes a robot, not a web page.

Have more questions? Join the waitlist and ask us.

Protect your release cadence with Testbed.

One bad build reaches every robot in minutes. Testbed sits on the OTA path you already run and checks a new build against the one it replaces before the fleet ever sees it.

  • Shadow mode first. A dry run on your own rollout telemetry.
  • Works with the OTA path you already run.
  • We see your metadata, not your data.

Join the waitlist