Failure modes, not failures: clustering what breaks

A list of failed episodes is a to-do list. A list of failure modes is a strategy.

Failure modes, not failures: clustering what breaks

After an evaluation run, most teams have a list: episode 4,812 failed, episode 4,830 failed, and so on. Someone watches them. Someone forms a theory. The theory is written in a document that a new team member will not find. This is how robotics programs relearn the same lesson every quarter.

From instances to classes

A failure mode is a class of failure with a shared cause. "Occlusion at grasp" is a failure mode. "Late-afternoon glare on the east-facing shelves" is a failure mode. Episode 4,812 is an instance of one of them. Fixing the instance — relabel this episode, retrain — does nothing for the class. Fixing the class fixes every instance you have and every instance you have not yet recorded.

How agents cluster

Agentuor's agent-intelligence layer indexes failed episodes by the conditions in their metadata, the geometry of their labels, the moment of failure in the timeline, and the visual and point-cloud features around that moment. It groups them, names the groups by the conditions they share, and presents each with representative examples and a count.

The names are proposals. A reviewer confirms, renames, splits, or merges them. Once confirmed, a failure mode becomes a first-class object in the workspace: it has an owner, a status, a linked evaluation suite, and a set of recommended actions.

You cannot manage what you have not named.

From cluster to action

Each failure mode carries recommendations in the same vocabulary as the rest of the platform. For occlusion at grasp: collect sixty episodes with partial occlusion of the target, add the cluster to the hard evaluation suite, clarify the guideline on when an occluded object should still be tracked. For glare: add lighting metadata to collection, record two afternoon sessions on the east shelves, slice evaluation by time of day. The recommendations become tasks in the workspace, and the failure mode's status tracks them.

Watching the cluster over time

Because failure modes persist, you can watch them. Did the occlusion cluster shrink after the collection run? Did a new cluster appear when the fleet moved to a new facility? Is the glare cluster seasonal? These are questions that were unanswerable when failures were rows in a spreadsheet, and they are the questions that decide whether a deployment is ready.

The shift from failures to failure modes is small in tooling and large in outcome. It turns evaluation from a report into a plan.

Published February 24, 2026 · All posts

← OlderLiDAR annotation for robotics: what image tools get wrongNewer →One schema from research lab to production fleet

Working on the same problems?

We'd like to see your data.