Pick the work, not the model
A model is the least differentiated part of an AI workflow. The expensive parts are everything around it: getting clean input in, defining what counts as correct, and handling the cases the model gets wrong.
Start from the queue. Look for work that arrives often, arrives as text or documents, follows rules a competent colleague could apply, and produces an output somebody else has to consume. That combination returns value fastest.
Be suspicious of workflows where success is a matter of taste. If two experienced people would disagree about the right answer, the model is not the constraint. The definition of the job is.
- Frequent enough to produce a signal within weeks
- Input is mostly text, documents or semi-structured records
- A competent colleague could state the acceptance criteria
- Failures are recoverable rather than final
If you cannot write the acceptance test before building, you will not be able to tell afterwards whether the model works.
Normalise the intake before you involve the model
Most disappointing AI features are retrieval and intake problems wearing a model costume. A workflow fed by forwarded email, shared drives and inconsistent spreadsheets spends its accuracy budget on formatting.
Front-load the deterministic work. One capture path, a required-fields contract, file type and size limits, deduplication on arrival, and a document store that can be queried by metadata. Complicated intake makes every later component harder to evaluate, because you can no longer separate a bad extraction from a bad source.
This is unglamorous plumbing, and it is also the part that keeps improving after launch, whereas model choice tends to converge quickly.
- A single intake endpoint, even when the senders are people
- Schema validation at the boundary with a visible rejection path
- Content hash on arrival so duplicates are never processed twice
- Metadata that reflects how the business already talks about the record
Ground the output on a source the business already trusts
Retrieval quality decides answer quality. Models are good at phrasing and unreliable about what they have actually seen, so the design question is how to get the right passage in front of the model and how to prove afterwards that it was the right one.
Chunk on meaning rather than on a fixed character count. Carry business metadata through the pipeline so a query can be filtered by document type, entity, date or permission before ranking happens. Attach a citation to every generated claim, and allow a confident refusal when the retrieved set does not support an answer.
Hybrid retrieval matters more than it first appears. Exact identifiers, clause numbers and product codes behave badly as embeddings, and those lookups are often the ones the workflow depends on.
- Hybrid keyword and vector retrieval, because exact identifiers fail embeddings
- Permission filters applied before ranking, not after generation
- Citations attached to individual claims rather than tacked on at the end
- An evaluation set drawn from real historical questions
retrieval:
index: policy_and_contract_docs
strategy: hybrid
keyword_weight: 0.4
filters: [entity_id, doc_type, effective_date]
require_citation: true
on_insufficient_evidence: refuse
Design the failure path before the happy path
A workflow that only works when everything succeeds is a demonstration. Real workflows meet malformed input, timeouts, partial vendor responses and rate limits, and each of those needs a defined outcome rather than an exception log.
Wrap every model call and every external action in a policy: a timeout, a retry budget with backoff, and an idempotency key so a retry cannot create a second order. Below a confidence threshold, the workflow routes to review instead of guessing.
Long-running work belongs on a queue with a dead-letter path, because a request-scoped process that dies takes its decision with it. Persist the intermediate state so a failed run can be inspected and resumed rather than re-inferred from logs.
- Timeouts on every outbound call, including the model endpoint
- Retry budgets per workflow so retries cannot become a load amplifier
- Idempotency keys on any action that creates or moves value
- Dead-letter handling with a named owner and a replay path
policy:
model_timeout_ms: 20000
retries: 3
backoff: exponential_jitter
idempotency_key: "workflow:{run_id}:{step}"
min_confidence: 0.82
on_failure: dead_letter_and_notify
Decide what a person signs, and what a person merely reviews
Some steps should never run unattended: anything that commits money, changes a contractual term, or notifies a regulator. Other steps are safe to automate with sampling, because the cost of an error is small and recoverable.
Approval only means something if the reviewer is given the evidence and the choice. Show the source, the extracted fields with confidence, the rule that fired, and what happens on each side of the decision. A queue nobody works is a system that has quietly stopped.
Give the queue an owner and a service expectation. If review is the bottleneck, either the thresholds are too permissive or the workflow is automating the wrong thing.
- Reserve human sign-off for consequential, hard-to-reverse actions
- Automate the reversible, the low-value and the easily detected
- Give every approval queue an owner and a review expectation
- Sample automated decisions so quality is measured rather than assumed
Measure the workflow, not the model
Model-level metrics are easy to collect and hard to act on. A self-rated answer quality tells nobody whether the claims team got its cases resolved faster.
Measure the operational shape instead: time from arrival to decision, first-pass accuracy against the acceptance criteria, escalation rate to review, and cost per completed item. These are numbers a business can hold you to, and they respond to changes in intake and retrieval as much as to changes in the model.
Keep an evaluation set of real historical cases and run it on every prompt, retrieval or model change. Treat a drop as a build failure rather than something to notice in production.
- Cycle time from input to usable output
- First-pass accuracy against the stated acceptance criteria
- Escalation rate, and which triggers are driving it
- Cost per completed item, tracked per workflow
A prompt change that improves the demo and drops the evaluation score is a regression.
Takeaways
- Choose the workflow before choosing the model
- Fix intake and retrieval before tuning generation
- Give every outbound call a timeout, a retry budget and an idempotency key
- Route low-confidence output to review instead of guessing
- Optimise cycle time and escalation rate, not model benchmarks