The review queue is the product
Once a person is in the loop, the workflow’s capacity is the reviewer’s capacity. An approval screen that takes longer to use than the task it replaces has saved nobody anything, however good the extraction underneath it is.
That makes queue design the central engineering problem rather than a detail. How many items a reviewer sees, in what order, with how much context, decides whether the human step is a real control or a rubber stamp with a deadline attached.
Design the review surface alongside the automation, from the first version. Retrofitting it later means interrupting a workflow that people have already learned to distrust.
Measure a reviewer’s time per item before optimising the model’s accuracy.
What a reviewer actually needs to see
A reviewer cannot verify a value they cannot trace. Show the source document, the extracted field beside the passage it came from, the confidence, and the rule that fired. When the two disagree, the reviewer should see both without opening another system.
Make the decision cheap. Keyboard-first, one action to accept, one to reject with a reason, one to escalate. Rejecting without a reason makes the correction data useless, and useless correction data is why quality stops improving.
Show the consequence of each option where there is one. For a payment or a contract change, the amount and the counterparty belong on the same screen as the approve button.
- Source document and originating passage visible beside the field
- Confidence score and the rule that produced the decision
- Keyboard-complete accept, reject-with-reason and escalate paths
- Consequence of approval shown next to the approval control
review_item:
field: policy_expiry
extracted: 2026-03-31
confidence: 0.61
source: { document: renewal_terms.pdf, page: 4, line: 112 }
rule_failed: date_within_notice_window
actions: [accept, reject_with_reason, escalate]
Order the queue by risk, not by arrival
First-in-first-out treats a trivial field correction exactly like a rejected payment. Order by the combination of value at risk, likelihood of being wrong, and how long the item has already waited.
Batch similar items so a reviewer can compare several documents against each other, which catches systematic errors that isolated review misses. Cap the batch so it still fits in one sitting.
Give the queue a service expectation and an escalation path when it is breached. A queue that grows quietly is how a team discovers months later that the automation has been failing and nobody was looking at the output.
- Priority derived from value at risk rather than arrival time
- Batched comparison for related document types
- A defined service expectation with escalation when breached
- Aged-item alerting so a growing queue cannot go unnoticed
Automation decays as the world moves
A review queue is not a permanent feature. As extraction improves and thresholds tighten, correct items should stop arriving at all. If queue volume never falls, either the thresholds are too loose or the workflow is not learning from corrections.
The subtler problem is reviewer drift. Over time people approve faster, checks become habitual, and the control weakens with no visible change. Sampling approved items for a second review is the cheapest defence, and it produces a measured accuracy figure instead of an assumption.
Rotate difficult items between reviewers. One person handling every ambiguous case concentrates the knowledge and the risk in the same place, which is a poor property for a control.
- Queue volume treated as a health metric that should trend down
- Sampled second review of automated approvals
- Rejection reasons captured in a structured form, not free text
- Rotation of hard cases across reviewers
- Thresholds revisited when source document formats change
Close the loop from corrections to evaluation
Reviewer decisions are the most valuable data the workflow produces, and most systems discard them. Route every rejection and every overridden suggestion into an evaluation set, then re-run that set on every prompt, retrieval or model change.
Without the loop, model changes are judged by opinion. With it, a change that improves the obvious cases and regresses the rare ones is caught before production, which is the difference between tuning and gambling.
Use the labels you already have rather than collecting more. The cases reviewers reject most often are usually the ones worth a second look, and they are free.
- Rejections and overrides captured as labelled examples
- Evaluation set run on every prompt, retrieval and model change
- Rare and high-value cases prioritised over raw volume
- Reviewer agreement tracked as a quality signal in its own right
- Corrections fed into retrieval, not only into fine-tuning
Know when not to ask a human
Approval on every item trains reviewers to approve everything, which turns the control into theatre. Some fields are low value, low consequence and cheaply detected, and gating them only consumes the attention you need for the fields that matter.
Automate those with sampling: apply automatically, review a small fraction, and track the error rate. If the sampled error rate exceeds what the consequence justifies, tighten the threshold and put the field back in the queue.
Define explicitly which decisions are never delegable, whether to a model or to a reviewer override, and enforce it in code. A policy that lives only in a document erodes.
- Automate low-consequence fields under a sampled audit
- Escalation as the default model rather than approval
- Thresholds set per field rather than per workflow
- A written list of decisions that are never delegable
A control that checks everything equally will be checked thoroughly for nothing and carelessly for everything.
Takeaways
- Capacity, not accuracy, is the constraint once a human is in the loop
- Show the source, the confidence and the rule beside the field
- Order the queue by risk and give it a service expectation
- Turn rejections into evaluation data or stop collecting them
- Automate low-consequence fields with sampling rather than gating everything