Observability, Drift, and Maintenance
Why this matters
A system that cannot be inspected will eventually fail in ways you do not notice until the damage is done. The dangerous failures are not crashes. They are quiet: the agent that slowly starts classifying a subset of cases wrong, the constraint that silently stopped applying after a model update, the response that has been going out slightly off-brand for three weeks. Launch without observability is just deferred incident response.
What you will be able to decide after this
- Whether the solution's behavior is inspectable on demand — not just after a complaint.
- What the early-warning signals are, and who is responsible for seeing them.
- How a change gets made without losing control of what the system does.
Core lesson
Design for five things:
- Logs of what the agent saw, decided, and did. The decision trail: the input it received, the output it produced, and the route in between. If a case goes wrong, the log is how you answer "what did it see, and why did it decide that?"
- Visibility into tool calls and outputs. What the agent invoked, with what arguments, and what came back. An agent that calls tools invisibly is an agent you cannot audit.
- Detection of drift — performance or behavior changing over time. Drift is normal; models, inputs, and people all change. The failure is not detecting it until it matters. A subset accuracy number, a rejection rate, a complaint count — a signal that moves when behavior moves.
- Clear ownership of monitoring and response. A named person (or role) watches the signals, and a named response when a signal trips. "The team will keep an eye on it" is not ownership.
- A plan for updating constraints, prompts, or models without losing control. Every change goes through the same gate: test against known cases, check the constraints still hold, release with a rollback. The system that cannot be updated carefully will be updated carelessly, or not at all.
The observability and maintenance plan has five sections, and each must be fillable:
- What will be logged — the decision trail and tool calls, at a level that answers "what did it see and why did it decide that?"
- What signals indicate trouble — the named numbers that move when behavior moves.
- Who watches the signals — a named role, not a team in general.
- How often the system is reviewed — a cadence tied to how fast the failure would hurt.
- How changes to the system are tested and released — the gate every constraint, prompt, or model change passes through, and the rollback.
Worked example
A support-triage solution has been running for two months. The logs show every case the classifier saw, the class it assigned, and the draft it produced. The review signal is simple: a weekly check of classification accuracy on escalated cases, which the case team lead owns. In week six, the accuracy on one subset — new product line inquiries — drops from 92% to 81% in a week. Nothing else moved. The log shows what changed: the new product's email templates use a term the agent was never shown, so a chunk of cases now classifies to the wrong tier and the drafts read wrong. The lead sees the number before a customer does, the drafts are rerouted, and a constraint update (five new template terms, tested against the last two months of replayed cases) goes out through the change gate: replayed, checked, released, rollback one line. The quiet failure surfaced at 81%, not at the customer. That is the entire point of the plan.
Practice
For your solution, write the trouble signals as numbers — not adjectives. Then write who watches each one, by name or role, and the cadence. Mark any signal that nobody owns.
Apply — produce the artifact
Write a short observability and maintenance plan for the solution: what will be logged, what signals indicate trouble, who watches the signals, how often the system is reviewed, and how changes to the system are tested and released.
Verify
Simulate a quiet failure — for example, the agent starts returning plausible but incorrect results on a subset of cases. Confirm that your plan would surface it before major damage occurs. If not, strengthen the signals and ownership.
Sources
This module is original practice guidance based on the authoring standard and does not depend on a specific external factual claim. Editorial review is still required.