Workflow Automation

Automating operational workflows

Most operational automation needs no model at all. Idempotency, explicit states, exception queues and scheduled reconciliation do the bulk of the work.

Note 7 min read All insights

Handoffs are where work is lost

Operational failures are rarely caused by a difficult step. They happen in the transfers: the person who retypes a value, the spreadsheet nobody reconciles, the queue that grows because nothing watches it. Model the transfers first.

Describe the process as explicit states with named transitions, and record who or what is allowed to make each one. A process that exists only as a habit in three people’s heads cannot be automated safely, because there is nothing to enforce.

Record every transition with the actor, the time and the originating state. That record is what makes a bad outcome diagnosable and what makes a replay safe.

workflow: supplier_invoice
states: [received, extracted, matched, approved, posted, paid]
transitions:
  received -> extracted: { actor: pipeline }
  extracted -> matched:  { actor: rules, on: confidence_ok }
  matched  -> approved: { actor: human }
  approved -> posted:   { actor: pipeline, idempotent: true }

If a state exists only in someone’s understanding, it will eventually be skipped.

Make every step replayable

Queues deliver at least once and networks retry, so every handler will eventually run twice. Design for that from the start with an idempotency key derived from business identity rather than from the request, and store the outcome so a repeat can return the same answer.

Key on the natural identity of the thing being done: the invoice number, the document hash, the enquiry identifier. A synthetic event identifier makes a handler idempotent only for that delivery, not for the underlying intent.

Persist intermediate state rather than holding a workflow in memory. It makes long-running processes resumable after a deploy, and it makes the current position of any item queryable, which is worth more than the elegance of a single process.

  • Idempotency keys derived from business identity
  • Stored outcomes so repeats return rather than re-execute
  • Checkpointed progress for long-running workflows
  • State committed before a side effect is announced
handler: post_supplier_invoice
idempotency_key: "invoice:{supplier_ref}:{invoice_number}:{revision}"
on_duplicate: return_stored_result
transaction: { commit_state_before_side_effect: true }

Automate the rules deterministically, use models for judgement

Most operational automation needs no model at all. Eligibility, pricing, thresholds, approval limits and reconciliation are rules, and rules belong in code where they are testable, fast and reviewable.

A model earns its place where the difficulty is reading ambiguous input: classifying an incoming document, extracting fields from something never designed to be read, summarising a long history, drafting a first response. In each case the output feeds a deterministic process rather than replacing it.

Keeping the two separate makes the system easier to reason about. The deterministic part is where accountability lives, and it should be readable by someone who does not evaluate models.

  • Rules engine for eligibility, pricing, limits and thresholds
  • Model use confined to reading, classifying, extracting and drafting
  • Deterministic validation applied to every model output
  • Escalation paths defined before launch rather than discovered during it

Exception handling is the real workload

Every automated process produces a small volume of work that cannot be automated, and that volume decides whether the project succeeds. Design the exception path with more care than the happy path.

Give the queue an owner, a service expectation and tooling to act on it. Someone must be able to inspect an item, see what the system attempted, fix or reject it, and replay it, without a developer and a database session.

Bulk repair matters more than individual handling. Most exceptions come from one upstream change affecting many items, so the useful capability is a controlled re-run over a filtered set rather than a queue a person empties one entry at a time.

  • Every dead-letter queue has a named owner
  • A review and replay interface rather than a log file
  • Alerting on queue depth and oldest item age
  • Bulk re-run over a filtered selection of failed items
  • A scheduled reconciliation comparing state across systems

Reconcile rather than trust

Two systems that exchange messages will eventually disagree. The question is whether you find out from the customer or from a comparison you run yourself. Reconciliation is the mechanism, and it should be routine rather than an exercise performed during an audit.

Compare the states that matter, define a tolerance, and alert when the difference crosses it. A reconciliation reporting a large constant discrepancy is worse than none, because people learn to ignore it. Investigate the systematic part until the residual is genuinely small.

Keep the report as the record. When someone asks what happened to a specific transaction, the answer should be a query rather than an investigation.

  • Scheduled comparison of source and destination state
  • A stated tolerance, with the systematic difference investigated
  • Alerting on the residual rather than on the total
  • A per-item audit trail linking both sides of every movement

Ship it beside the old path, then remove the old path

The safest first release of an operational automation runs alongside what exists. Process items through both paths, compare the outcomes, and serve the new result only once you trust the disagreement rate. That costs compute and buys most of the certainty.

Cut over one capability at a time rather than one workflow at a time, so a failure stays small and understood. Keep the previous path runnable for long enough that it is a real option rather than a theoretical one.

Then remove it. Dual-running is the most common reason operational automation costs more than it saves, and removal is the step that gets deferred indefinitely.

  • Shadow mode against real traffic before serving results
  • Cutover per capability, with the old path still runnable
  • A decommissioning date agreed at launch
  • A single, documented way to reverse a cutover

Automation that cannot be turned off is a permanent change to how the business operates. Be deliberate about that.


Takeaways

  • Model states and transitions before automating anything
  • Key idempotency on business identity, not on the request
  • Keep rules deterministic and use models only to read
  • Treat the exception queue as the workload to design for
  • Reconcile on a schedule, and commit to removing the old path

Facing something similar?

These articles describe how we think about a problem. A conversation is where we would apply it to yours, against your actual constraints.

Start a conversation All insights