Introducing the agentic feedback loop

Agents that watch annotation as it happens, fold in reviewer decisions, and hand every annotator the context they were missing.

Introducing the agentic feedback loop

Today we are describing the part of Agentuor we have spent the most time on: the agentic feedback loop. It is the reason the platform exists, and it is easiest to explain by following one label.

Following one label

An annotator opens a manipulation episode. As they draw a cuboid around a tote, an agent compares the geometry with the LiDAR extent, with similar frames elsewhere in the dataset, and with the guideline currently in force. It proposes a small height adjustment and shows why: the point cloud extends four centimeters above the drawn box. The annotator accepts.

Later, a second agent notices that the wrist-camera projection of that cuboid disagrees with the head-camera view. It flags the label for review — not with a red icon, but with a specific note about which views disagree and by how much.

A reviewer adjudicates. They keep the annotator's box and add a guideline clarification: measure totes to the rim, not the lid. That clarification is pushed to the forty-two open tasks containing the same object class. The next annotator to draw a tote sees it before they start.

Five stages, one loop

The loop has five stages: annotate, analyze, recommend, review, improve. What makes it a loop rather than a pipeline is that the last stage feeds the first. Reviewer decisions calibrate future recommendations. Clarifications reach open work. Evaluation failures become collection tasks. Quality compounds instead of resetting with each batch.

The dataset carries the memory of every decision, so the team does not have to.

Specific actions, not scores

We spent a long time on the language of recommendations. A confidence score alone is a poor recommendation: it tells the annotator that something might be wrong without telling them what to do. Our agents recommend actions in the vocabulary of the guideline — split this event at frame 1,248, add a boundary where the object leaves the frame, mark the outcome, send this to expert review — each with confidence and evidence attached.

Humans decide, always

No recommendation is applied automatically. Pre-labels can be shown; nothing enters the dataset until a person confirms it. This is not a limitation we intend to relax as the agents improve. It is the design. A dataset in which every entry is a human decision, with a recorded reason, is a dataset you can audit, defend, and trust.

The feedback loop is live for early partners today across 2D, 3D, LiDAR, and video workflows. We will share what we learn as it meets more data.

Published September 16, 2025 · All posts

← OlderThe anatomy of a robot episodeNewer →Confidence is not enough: why every recommendation needs provenance

Working on the same problems?

We'd like to see your data.