Now onboarding robotics teams

Collect, annotate, and ship
robot-ready data.

Agentuor is data infrastructure for Physical AI. Agents watch annotation as it happens and turn every reviewer decision into a better dataset.

Annotate Analyze Recommend Review Improve agentic loop
the agentic feedback loop

Agents that watch the work, not just the output.

Agentuor agents analyze annotation as it happens, fold in reviewer feedback, and turn every correction into an insight the next annotator — and the next model — can use.

1

Analyze in real time

Every stroke, box, and event marker is checked against the guideline and the rest of the dataset while the annotator is still on the frame.

2

Recommend with provenance

Suggested labels and workflow actions carry a confidence score and a traceable reason. Nothing is applied without a person confirming it.

3

Improve continuously

Reviewer decisions retrain the recommendations, tighten guidelines, and surface the failure modes worth fixing next.

gripper · 0.97 tote · 0.91 occlusion · 0.78 Suggested next stepSplit event at frame 1,248
next-step guidance

Specific actions, not vague scores.

Instead of a red flag, annotators get a concrete suggestion they can accept, edit, or reject — each one logged with who decided and why.

Split this eventTwo grasp attempts detected inside one "pick" segment.
0.94
Add a boundaryObject leaves the frame at 00:41 — end the track here.
0.89
Mark the outcomeEpisode ended without a success/failure label.
0.97
Send to expert reviewAmbiguous occlusion — matches 14 similar disputed cases.
0.81

What makes Agentuor stand out

Humans decide, always

Agentuor is built so that every label is a human decision. Agents propose; a person confirms. No auto-apply, no silent changes — a dataset you can audit and defend.

Recommendations in real time

Suggestions arrive while the annotator is still on the frame — a boundary to add, an event to split, an outcome to mark — each with confidence and the evidence behind it.

Truly multimodal

2D, 3D, LiDAR, and video on one synchronized timeline with joint states and force-torque. Many tools bolt point clouds onto an image editor. We started from the geometry.

Enterprise-grade governance

Role-based access, complete audit trails, region choice, and hybrid deployments that keep raw data in your storage — from the first research dataset onward.

What can Agentuor do for you?

Provenance on every label

See what evidence produced a suggestion, who accepted or changed it, and which guideline version applied. Audit trails export with the data and attach to model cards.

Quality that compounds

Automated validation catches the obvious. Experts see only the cases that need judgement. Every decision recalibrates the next recommendation, so quality rises with each batch.

Coverage you can measure

Agents chart the corpus by lighting, clutter, and object class, deprioritize redundant footage, and turn blind spots into the next collection session.

Feel the simplicity

You don't need a data-engineering team to get going. Connect a robot, a rig, a simulator, or a bucket through our SDK; episodes are segmented and described automatically. We give you room to run collection, annotation, review, and evaluation in one workspace — with three delivery models available on the same dataset.

Connects to:
Robots & rigsSimulatorsCloud storageTraining pipelinesREST API
inside the workspace

Integrate, watch, act.

Plug it in

Push episodes with a few lines of code

await agentuor.ingest({
  project: "warehouse-pick",
  episode: recording,
  metadata: { site: "seville-1", shift: "pm" }
});
// agents begin profiling immediately
10 min setupREST APIPython SDK
Episode 4,812 · graspLive
t = 00:41.20 · suggested: split event
Regrasp detected inside "pick" — two attempts merged. Split at frame 1,248?agent · confidence 0.94 · evidence: gripper reopen + force dip
t = 00:43.05 · outcome missing

QUALITY

Agreement8.8
Geometry9.2
Continuity8.1

FLAGS

Merged eventFlagged
Track breakClear
OutcomeMissing

Act when it matters

Ambiguity routed with full context

Review requestedAmbiguous occlusion

Wrist-camera projection disagrees with head-camera view by 6 cm. Matches 14 similar disputed cases in this dataset.

Episodeep_4812
Guidelinev7 · rule 4.2
Confidence0.81
  • Routed to a perception lead
  • Decision recorded with reason
  • Clarification reaches 42 open tasks

One schema, lab to fleet

The dataset a research team starts with is the dataset the production fleet runs on. Growth adds rows, never rebuilds tables.

Find the patterns hiding in millions of frames

Agents cluster ambiguous examples and recurring failures by cause, so you fix the class of problem — not one instance at a time.

Guidance in the annotator's language

Not "confidence 0.6" but "split at frame 1,248". Recommendations speak the vocabulary of your guideline, with evidence attached.

Your team, ours, or both

Keep sensitive work in-house and burst volume to our managed workforce. One guideline, one quality standard, one provenance trail.

the full data journey

One platform from first recording to fielded robot.

01

Data collection

Teleop, sim, and fleet logs ingested with synchronized sensors and episode metadata.

Learn more
02

Multimodal annotation

2D, 3D, LiDAR, and video workflows in one Annotation Studio.

Learn more
03

Quality & review

Automated validation plus expert review, scaled by agents that route the right cases.

Learn more
04

Model evaluation

Edge-case and real-world condition testing to harden robustness before deployment.

Learn more
05

Agent intelligence

Pattern and failure-mode discovery across your entire dataset.

Learn more
built for robotics data

Everything multimodal data needs to become training data.

Multimodal annotation

Bounding boxes, cuboids, segmentation, keypoints, event timelines, and LiDAR point labels — synchronized across every sensor on the robot.

Learn more

Pattern and failure-mode discovery

Agents cluster ambiguous examples and recurring mistakes across millions of frames so you fix the class of problem, not one instance.

Learn more

Expert review at scale

Automated validation catches the obvious; domain experts get only the cases that need judgement, with full context attached.

Learn more

Real-world model evaluation

Slice performance by lighting, clutter, occlusion, and embodiment. Know where a policy breaks before the robot does.

Learn more

Enterprise-grade governance

Role-based access, audit logs, data residency options, and controls built for regulated and safety-critical programs.

Learn more

Research to production

Start with a single research dataset and grow into fleet-scale pipelines without changing tools or schemas.

Learn more
5stages of the data journey, one platform
4annotation modalities: 2D, 3D, LiDAR, video
3delivery models: SaaS, managed, hybrid
100%of recommendations carry confidence and provenance
why physical ai is different

Robot data is long, multimodal, and unforgiving.

An image dataset is a pile of pictures. A robotics dataset is minutes of synchronized sensors, dozens of joints, ambiguous outcomes — and a bad label shows up as a robot doing something it shouldn't. Agentuor was built for that reality, not adapted to it.

Head RGBWrist RGB-DLiDARJoint statesForce-torque t = 0.00 one clock, every sensor
synchronized by default

Every stream on one clock.

Head and wrist cameras, LiDAR, joint encoders, force-torque — aligned at ingest with drift correction, so a boundary placed on the video is the same boundary in the point cloud and the telemetry.

  • Dropped frames and clock skew flagged automatically
  • Episodes segmented from continuous recordings
  • Metadata inferred with confidence, confirmed by an operator
About data collection
coverage you can see

Stop labeling what you already have.

Agents cluster the corpus by lighting, clutter, object class, and motion, then chart it. Redundant footage is deprioritized; blind spots become the next collection session. Annotation budget goes where the robot is weakest.

About agent intelligence
Scenario coveragelighting × clutter
dim → brightsparse → cluttered

Faint cells are the scenarios your robot hasn't seen enough of. Agents point collection there next.

provenance on everything

If a label is in the dataset, you can see why.

Every recommendation, edit, flag, and decision is recorded with who made it and what evidence they had. Audit trails export with the data and attach to model cards.

Provenance · label #48,213guideline v7
  1. Pre-label proposedagentcuboid "tote", confidence 0.91, evidence: 3 similar frames + LiDAR extent
  2. Accepted with editannotatorheight adjusted +4 cm to match point cloud
  3. Flagged for reviewagentdisagreement with wrist-camera projection
  4. Adjudicatedexpertkept edited cuboid; guideline clarified: "measure to rim, not lid"
  5. Clarification pushedsystemreaches 42 open tasks with the same object class
Occlusion at grasp1,204 episodes Glare, late afternoon388 episodes Track identity swap Missing outcome Recommended fixCollect 60 occluded grasps

Left: the life of one label. Right: failure modes grouped by cause, each with a recommended fix.

the loop closes

Evaluation talks back to collection.

A failing slice in evaluation links to the episodes and labels behind it — and becomes a collection or relabeling task in the same workspace. Nothing is lost between tools, because there is only one.

CollectAnnotateReviewEvaluateImprove what evaluation learns, collection hears
what changes

Before and after one platform.

beforeFive tools, zero memory

  • Recorder, labeling vendor, QA spreadsheet, eval notebook, ticket queue
  • Boundaries drift between annotators and across weeks
  • Reviewers see labels, not the reasons behind them
  • Evaluation reports an average; nobody knows which condition failed
  • The same ambiguity is re-argued on every batch

afterOne workspace, compounding quality

  • Collection, annotation, review, evaluation under one schema
  • Agents propose boundaries from state; people confirm
  • Every decision carries confidence, evidence, and a name
  • Performance by slice, traced to source episodes
  • Ambiguities resolved once and pushed to every open task
how we build

Four principles behind every feature.

01

Humans decide

Agents recommend. A person confirms. No label enters the dataset any other way.

02

Provenance is not optional

Confidence, evidence, and a name on every decision — exportable, auditable, permanent.

03

One schema, lab to fleet

The research dataset and the production pipeline never need a migration.

04

Quiet software

The best data tool is the one annotators stop noticing.

quick answers

Frequently asked

Do agents ever change my data without approval?

No. Recommendations are never applied automatically. Every change is confirmed by a person and logged with a reason.

Which data types are supported?

2D images, 3D data, LiDAR point clouds, and video, plus synchronized telemetry such as joint states and force-torque readings.

Can I start in SaaS and add managed capacity later?

Yes — it's the most common path. Ontologies, guidelines, and quality history carry over unchanged.

Where is my data stored?

In the region you choose, or in your own storage under a hybrid deployment.

All questions

two ways to run it

Your workspace, or ours — or both.

Self-serve SaaS

For teams who want full in-workspace control

  • Your annotators, your guidelines, your queue
  • Agent recommendations inside every task
  • Dashboards, exports, and API access
About SaaS

Managed services

For teams who want end-to-end execution

  • We staff, train, and run the annotation workforce
  • Domain expert review included
  • Delivery against agreed quality SLAs
About managed

Hybrid

For teams who want to mix both

  • Keep sensitive or expert tasks in-house
  • Burst volume to our managed team
  • One dataset, one quality standard
About hybrid
ready when you are

Bring your first dataset. We'll show you what the agents find.

A 30-minute walkthrough on your own data, or on a sample robotics dataset if yours isn't ready yet.