Skip to content
Dx
05Leadership & Adoption · 5 of 5

Ownership When It Breaks

~60 min of real work

Ownership When It Breaks

Why this matters

Every system fails. The organizational question is whether ownership is clear when it does. Diffuse ownership produces slow, incomplete responses and repeated failures: the monitoring is "the team's" (which is to say no one's), the shutdown requires a meeting, the communication waits for a decision, and the learning never happens because nobody owns the after-action. The failure that is owned is an incident. The failure that is not owned is a pattern.

What you will be able to decide after this

  • Who does what when a specific failure happens — without an argument at the moment.
  • Whether the system can be shut down quickly, and by whom.
  • Whether failures produce learning, or just repeat.

Core lesson

Before any meaningful deployment, name four things:

  • Who is accountable for monitoring. The person who watches the signals — the drift numbers, the error rate, the anomalies — on a cadence. Monitoring is a job with a schedule, not a hope.
  • Who has authority to shut it down or roll it back. The named person who can stop the system immediately, no meeting required. If the answer is "we would convene and discuss," the system is not shut-down-able.
  • Who communicates with affected users or customers. The person who owns the message when something goes wrong — what happened, what is being done, what happens next.
  • Who owns the post-incident learning and change. The person who turns the incident into a constraint, a rule, or a checklist so the same failure becomes less likely.

The four roles can be held by the same person in a small organization. What they cannot be is unassigned. This maps directly to the accountability expectations in NIST AI RMF's Govern function and ISO/IEC 42001 — the frameworks say the same thing in committee language: someone has to be responsible, and it has to be specific.

Worked example

System: AI that drafts contract redlines for the legal team. Ownership map:

  • Monitoring: Legal Ops lead — weekly sample plus alert on confidence or volume anomalies.
  • Shutdown authority: Head of Legal — can disable the integration immediately.
  • Communication: Legal Ops plus the affected deal owner.
  • Learning: joint review by Legal Ops and the system owner within 5 business days of any significant incident.

The verification walk-through: a redline that misses a liability clause on a high-value deal. Monitoring catches it on the sample within the week — but the walk-through reveals the gap: the redline already went to the counterparty before the sample ran, and nobody had a rule for "any redline on a deal above the threshold waits for human review." The map is fixed before deployment: the threshold rule is added, and the deal-owner communication path is pre-named. The walk-through caught the failure the map would have missed.

Practice

For the system you will map, write the four roles as open slots and fill each with a name or role from your context. Mark any slot you filled with a department instead of a person, and decide who the person actually is.

Apply — produce the artifact

For one real or realistic AI system in (or near) production, write a short ownership map covering monitoring, shutdown authority, communication, and learning. Include names or roles, not just department labels.

Verify

Walk through a realistic failure scenario with the people named. If any critical action has no clear owner or requires a meeting to decide who decides, fix the map.

Sources

This module is original practice guidance based on the authoring standard and does not depend on a specific external factual claim. Editorial review is still required.