There is a paradox at the heart of assisted annotation. As recommendations improve, annotators learn that accepting them is usually right. Their attention drifts. Acceptance rates climb toward one hundred percent — including on the recommendations that are wrong. A ninety-five percent accurate agent with a one hundred percent acceptance rate produces a dataset with five percent silent errors, all of them wearing a human signature.
Naming the problem
Automation bias is well studied in aviation and medicine: the tendency to over-trust automated suggestions and under-apply one's own judgement. Annotation is a textbook environment for it — repetitive, high-volume, with a system that is right most of the time. Pretending it will not happen is not a strategy.
What we measure
For every recommendation type, we track its actual accuracy against expert adjudication and its acceptance rate per annotator. The gap between them is the signal. An annotator accepting ninety-eight percent of a recommendation type that adjudication shows to be eighty-five percent accurate is not being efficient. They are not looking.
We also measure time-to-decision. Acceptances that arrive faster than a person could plausibly have checked the evidence are flagged, not as misconduct, but as a calibration problem for that annotator on that recommendation type.
What we do about it
Evidence, not just confidence. As we wrote last year, showing the evidence behind a suggestion changes the question the annotator is answering, and acceptance rates fall toward true accuracy.
Seeded checks. A small fraction of recommendations shown to each annotator are deliberately wrong in ways the evidence reveals. Catching them confirms attention; missing them triggers a recalibration task.
Adaptive assistance. When an annotator's acceptance rate on a type detaches from its accuracy, assistance for that type is reduced for them — fewer pre-labels, more blank frames — until calibration returns.
Feedback to the person. Annotators see their own calibration. Most people, shown that they are rubber-stamping, stop.
The agent's accuracy is one number. The dataset's accuracy is the agent's accuracy filtered through a person's attention.
Why this matters for oversight
"Human in the loop" is only meaningful if the human is actually looking. A signature on a label that nobody examined is provenance in form but not in substance. Measuring automation bias — and acting on it — is how we make the human oversight we promise into something real rather than ceremonial.
We will publish more of what we learn as the platform meets more annotators and more data.