Skip to content
Dx
03Workflow & Systems · 5 of 5

Failure Handling at Workflow Level

~75 min of real work

Failure Handling at Workflow Level

Why this matters

Individual AI failures are expected. A model will produce a wrong answer; that is the known cost of the tool, and Core Judgment covers it. Workflow-level failures are the ones that quietly damage trust and results: a wrong recommendation advances, nobody is notified, the customer's case sits in a bad state, and the next run builds on the failure. A workflow that cannot surface and handle its own failures is not ready for production use.

What you will be able to decide after this

  • Whether the workflow's failure modes are visible, contained, owned, and recoverable.
  • Who is responsible when a specific failure happens — without an argument at the moment.
  • What must be recorded so the same failure becomes less likely.

Core lesson

Design for four properties:

  • Visible failure — you know when something went wrong. Detection is explicit: a check, a threshold, a rule, a person looking. If the only detection is "someone will probably notice," the failure is invisible by design.
  • Contained failure — it does not cascade silently. A wrong output stops at the checkpoint instead of flowing downstream, and downstream work does not build on unverified output.
  • Owned failure — a person is responsible for response. Not "the team," not "the system." A named role, with a named response.
  • Recoverable failure — there is a clear path back to a good state. The case can be replayed, the draft can be rejected and redone, the customer can be re-contacted. If the only path back is manual archaeology, it is not recoverable.

The simulation

The test of these properties is a simulation, run before the change ships. Take a realistic high-impact failure for the workflow you are redesigning. Example scenario:

The AI-assisted step produces a plausible but incorrect recommendation on a customer case that is already escalated. The human checkpoint is busy and the system is configured to auto-advance after 4 hours if no action is taken.

Walk the simulation:

  1. How does anyone know the recommendation was wrong? What detects it — a later check, a downstream person, the customer?
  2. Who is notified, and how fast? Name the person and the mechanism. "The case manager will see it" is not a notification.
  3. What is the immediate containment action? What stops the wrong output from advancing or reaching the customer, right now?
  4. How is the customer or downstream work protected? What happens to the case while the response runs?
  5. How does the workflow return to a correct state? The redo path — rejected draft, corrected record, replayed case.
  6. What gets recorded so this specific failure becomes less likely? The lesson goes somewhere the next run can use: a rule, a constraint, a checklist, a review.

If the answers are vague or ownership is unclear during the simulation, the design is not ready. The simulation is supposed to be uncomfortable. The failure of the simulation is the point of the simulation.

Worked example

In the simulation above, the team discovers three things. First: nothing detects the wrong recommendation until the customer replies, because the checkpoint only reviews what it sees and the auto-advance moved the case on. The fix: the auto-advance is removed — "No silent auto-advance" — and a missed review surfaces the case to the supervisor queue instead. Second: the immediate containment action is "reject the draft and re-queue the case," which already exists but was never named, so in the simulation nobody said it for six minutes. Third: the recording step is empty — there is nowhere a lesson lands. The brief now names the detection rule, the supervisor escalation, the reject-and-re-queue response, and a weekly review where failures are logged. The team runs the simulation twice more before the change ships. The design passes when no one in the room is guessing.

Practice

Write the six simulation questions for one high-impact failure in your own workflow — not the example. For each question, write the answer you would give today, and mark any answer that is a guess.

Apply — produce the artifact

For the redesigned workflow, write a short failure and recovery brief:

  • How major failure modes will be detected
  • Who is notified
  • What the immediate response is
  • How normal operation is restored
  • What gets recorded so the same failure is less likely next time

Then run one concrete failure simulation — like the example above or one drawn from your own work — and note what broke or felt unclear.

Verify

Conduct the simulation with the people who would actually be involved. If response or ownership is debated in the moment, revise the brief and the design before going further.

Sources

This module is original practice guidance based on the authoring standard and does not depend on a specific external factual claim. Editorial review is still required.