Skip to content
Dx
05Leadership & Adoption · 4 of 5

Pilots That Actually Teach

~60 min of real work

Pilots That Actually Teach

Why this matters

Many pilots are designed to confirm a decision that has already been made. The question is written so the answer is yes, the criteria are soft, and the end-of-pilot meeting produces an expansion regardless of what the numbers said. Useful pilots are designed to reduce uncertainty — and a pilot that cannot change the subsequent decision is theater.

What you will be able to decide after this

  • Whether the pilot's question is answerable — and whether the answer could change your mind.
  • What "success" and "failure" mean, decided before the data arrives.
  • Whether the pilot is scoped tightly enough to produce a real signal.

Core lesson

A pilot worth running has five parts, and each one must survive scrutiny:

  • A clear question it is trying to answer. One question, with a direction and a magnitude. "Does this help?" is not a question; "Does this reduce total handling time by at least 20% without increasing escalations?" is.
  • Success and failure criteria defined in advance. Written before the pilot starts, with numbers. Success and failure thresholds are the entire point: they are the contract that makes the end-of-pilot decision non-negotiable.
  • A limited scope and time box. Narrow enough to measure, short enough to end. The scope is the ticket category, the teams, the volume — named. The time box is the date when the decision gets made.
  • Real users or real workflow, not a lab demo. The pilot runs where the work actually happens, with the people who actually do it, under the actual conditions — including the bad days.
  • A plan for what will be measured and what decision will follow. The measurements named in advance, and the decision — expand, modify, or stop — attached to the thresholds.

Worked example

Question: "Does AI-assisted drafting reduce total handling time (including verification) by at least 20% on tier-2 support tickets without increasing customer escalations?" Success threshold: ≥20% net time reduction and no increase in escalations over 6 weeks. Failure threshold: <10% net reduction or any material rise in escalations. Scope: two teams, one ticket category, human review required on every reply. Decision at end: expand, modify, or stop. The design does the work: because the thresholds are written down, the week-four numbers — 12% net reduction, flat escalations — land between the thresholds, which forces the honest conversation: it is not a failure, and it is not a success; the pilot is modified (verification time on complex tickets is trimmed with a checklist) and extended rather than quietly declared a win. The thresholds made the ambiguity visible instead of letting it hide.

Practice

For the initiative you have in mind, write the primary question in one sentence with a direction and a magnitude. Then write what would convince you to expand it, and what would convince you to stop it — in numbers.

Apply — produce the artifact

Design a pilot for one initiative that has not yet been fully committed. State the primary question, the success/failure thresholds, the scope, the measurement plan, and the decision that will be made at the end.

Verify

Ask: "If the pilot fails against the criteria I wrote, will we actually stop or change course?" If the honest answer is no, the pilot is not designed to teach.

Sources

This module is original practice guidance based on the authoring standard and does not depend on a specific external factual claim. Editorial review is still required.