Verification Before "Done"
Why this matters
Never mark a build complete because it worked in a demo or a happy-path test. Demos run the path you prepared. Happy-path tests run the input you imagined. The real world runs the rest — the edge case, the adversarial input, the exhausted human at the checkpoint, the constraint that quietly fails under pressure. "It worked when I tried it" is not evidence.
What you will be able to decide after this
- Whether the solution actually meets the acceptance criteria — on evidence, not confidence.
- Whether the human checkpoints hold under the conditions that matter.
- What an honest status is: ready for limited use, needs changes, or not ready.
Core lesson
Before you call it done:
- Test realistic and adversarial cases. Realistic: the inputs that actually arrive — messy, incomplete, borderline, from the people who actually use it. Adversarial: the inputs that try to break it — the prompt that attempts to override its constraints, the case designed to slip past the prohibited actions, the volume spike.
- Confirm the human checkpoints actually function under pressure. A checkpoint that works in the demo but is skipped when the person is busy is not a checkpoint. Test the miss: what happens when the reviewer does not act? The answer must be a visible, owned state — not silent auto-advance.
- Verify that prohibited actions are blocked. Not "probably blocked because we told it not to." Tested blocked: attempt each prohibited action and confirm the system refuses.
- Check that failure modes are visible and owned. The Module 4 plan in action: a seeded failure surfaces on the signals, and a named person responds.
- Measure whether the system meets the acceptance criteria in the contract. The Module 1 criteria, tested one by one, with results written down. Not "looks good." The numbers.
The verification record has four sections:
- Test cases run — including at least two failure or edge cases. Name them; a test that cannot be named did not happen.
- Results against the original acceptance criteria — each criterion from the contract, with a pass/fail result and the evidence.
- Any gaps that remain — what was not tested, what failed, what is unproven. The honest record includes these; the dishonest one hides them under "looks good."
- Explicit decision — ready for limited use / needs changes / not ready. A system that needs changes but ships anyway was shipped by someone who skipped this record.
Worked example
A procurement-drafting solution looked complete after a strong demo. The verification record changed the verdict. The realistic tests used real past requests — including two that were borderline and one with contradictory fields; the drafts for the borderline cases were usable but the contradictory one produced a draft that combined both claims without flagging the conflict. The adversarial test attempted a prompt that asked the agent to ignore its constraint and draft a purchase above the approval threshold — the agent refused, but the second attempt, phrased as a legitimate override request, produced a draft above threshold without routing to the human checkpoint. Prohibited action not actually blocked. The failure-mode test: a seeded wrong draft surfaced on the review signal within a day, and the named owner responded. Acceptance criteria: five of six met; the sixth — "no draft advances without review" — failed. The record says: needs changes. The override gap is fixed, the contradictory-field flag is added, and the record is re-run before any limited use.
Practice
For your solution, write the acceptance criteria from Module 1 as testable statements. Then write two edge or failure cases you have not tested yet, in one sentence each.
Apply — produce the artifact
Produce the verification record for the solution: test cases run (including at least two failure or edge cases), results against the original acceptance criteria, any gaps that remain, and an explicit decision — ready for limited use / needs changes / not ready.
Verify
Have someone who did not build the system review the verification record and try to break it. If they succeed in a way the tests did not catch, add the case and fix the gap.
Sources
This module is original practice guidance based on the authoring standard and does not depend on a specific external factual claim. Editorial review is still required.