One loop, five stages, humans in charge.
Here is what happens to a robotics dataset inside Agentuor, from the first upload to the evaluation gate before deployment.
1. Ingest and describe
Data arrives from robots, teleoperation rigs, simulators, or cloud storage. Sensor streams are aligned on a common clock, dropped frames are flagged, and continuous recordings are cut into episodes. Agents infer metadata — task, embodiment, scene, outcome — and show it with confidence for an operator to confirm. Governance rules such as blurring and retention are applied here, and logged.
2. Annotate with assistance
Annotators open tasks in the Annotation Studio with the relevant guideline pinned and synchronized sensors loaded. Pre-labels and live recommendations appear inline: a suggested boundary, a merged event to split, an outcome that is missing. Each suggestion shows its confidence and evidence. The annotator accepts, edits, or rejects; nothing enters the dataset otherwise.
3. Validate and review
Every submission passes automated checks for schema, geometry, temporal continuity, and agreement. Agents rank the review queue by disagreement, ambiguity, and downstream impact. Experts adjudicate with side-by-side context and record their reasons. Recurring problems are tagged and clustered into failure modes.
4. Evaluate by condition
Held-out episodes are organized into versioned suites by lighting, clutter, occlusion, embodiment, and other conditions. Policy outputs are scored per slice. A failing slice links straight to the episodes and labels behind it, and agents propose the collection or relabeling most likely to fix it.
5. Improve continuously
Reviewer decisions calibrate future recommendations. Clarifications become guideline updates that reach every annotator. Evaluation findings become collection tasks. The dataset carries the memory of every decision, so quality compounds rather than resets with each batch.
Where people sit in the loop
- Annotators decide on every label and suggestion.
- Reviewers and experts adjudicate ambiguous cases and set guideline direction.
- Program leads configure what agents may recommend, to whom, and what requires sign-off.
- Administrators govern access, residency, and retention.