Skip to content
Dx
06Enablement · 5 of 5

Measure Whether Judgment Actually Improved

~75 min of real work

Measure Whether Judgment Actually Improved

Why this matters

The only meaningful measure of enablement is whether people make better decisions and produce better-verified work afterward. Everything else — the satisfaction scores, the completion rates, the session attendance — measures the event, not the outcome. It is possible to run an excellent-feeling program that changes nothing. The measurement exists to tell the difference.

What you will be able to decide after this

  • What the evidence of improved judgment actually is, for your group.
  • How to sample real work without creating a new bureaucracy.
  • Whether the enablement worked — on the work, not the mood.

Core lesson

Useful signals:

  • Quality of artifacts against the standards in Core Judgment — the constraints, the verification, the stop decisions, visible in the actual output.
  • Ability to apply the 10 Questions or net-value thinking to new cases — the filters used unprompted on work that was not covered in the training.
  • Reduction in low-value or high-risk AI use — fewer expensive or pointless runs, visible in what people actually do.
  • Clear ownership when something goes wrong — the post-incident note names a person and a lesson instead of blaming the tool.
  • Honest statements of uncertainty — "this part I still cannot verify" appearing where confidence used to be fake.

Satisfaction scores and completion rates are secondary at best. They are not worthless — a miserable program also fails — but they cannot stand alone as evidence. A high satisfaction score measures whether people enjoyed it. The artifact quality measures whether it worked.

The measurement approach has five parts, kept small on purpose:

  1. What you will look at — the artifacts or decisions: samples of drafts, new-task exercises, post-mortems.
  2. What "better" means — against the capability statement from Module 1, written as observable differences.
  3. How you will sample — real work from participants, a small set, on a schedule; not a special exercise built for the review.
  4. When you will review — a date, after enough real work has accumulated to tell.
  5. Who does the evaluation — someone who can apply the standards without being invested in the program's success.

Worked example

An enablement program for a drafting team ends with a 4.8/5 satisfaction score and 100% completion. The measurement approach was defined anyway: a sample of eight real drafts per participant — four from the month before the program, four from the month after — scored against the capability statement: constraints visible, verification performed, uncertainty stated. The score does not move on three of the six participants. Their drafts show the same pattern as before: fluent output, no constraint record, no verification trail. The satisfaction was real and the completion was real, and the judgment did not move. The finding is not "the program failed" — it is "the program produced attendance and warmth, and the design must change so the artifact evidence is part of the program itself." The next iteration builds the artifact into the session, as Module 2 of this pack requires, and the measurement runs again.

Practice

For your enablement effort, write what "better" would look like in the artifacts you can actually access — three observable differences, in one sentence each. Then mark which of the five signals you could actually see in your context today.

Apply — produce the artifact

Define a simple measurement approach for one enablement effort: what artifacts or decisions you will look at, what "better" means, how you will sample, and when you will review the results.

Verify

After the enablement activity, collect a small sample of real work from participants. Evaluate it against the capability statement from Module 1. If you cannot see a difference in judgment quality, the enablement did not work — regardless of how people felt about the sessions.

Sources

This module is original practice guidance based on the authoring standard and does not depend on a specific external factual claim. Editorial review is still required.