NEWProduction Web Themes & Turnkey ArchitecturesGet Lifetime Pass ($199) →
KNKomal Nakrani
Get All Access
ThemesDocsAll-Access PassGet All Access ($199)
Book overview
18/Applied AI Engineering

Diagnose Real Use and Incidents

Contain consequence, reconstruct layered evidence, test and disconfirm hypotheses, protect sensitive traces, verify recovery, and change durable artifacts.

The metric improves while the product fails

Patchwork’s bounded text cohort shows higher engagement. Users click more candidate cards. The service is green. Latency stays within the text envelope.

Two adjudicated reports say recommended parts did not fit. Both involve long-tail seller-authored records. One record contains instruction-like language that escaped the known-token quarantine. The ranking path treated a seller claim as an attractive feature.

The first hypothesis is the model regressed. It is plausible and wrong. The model adapter did not change. Paired replay fails only when the corrupted seller field is present. The evidence supports interacting data-context, product-metric, and control conditions:

  • untrusted seller content entered a feature path;
  • the token-pattern control missed novel phrasing;
  • engagement rewarded attractive unsupported results;
  • aggregate service health hid the behavior failure.

The correct response is not to tune ranking while exposure continues. Stop the cohort, restore deterministic search, quarantine affected records, preserve privacy-conscious evidence, test hypotheses, correct the boundary, replay evaluation and controls, and route any renewed exposure to release authority.

This chapter advances PF-11 to v0.2 with a fully fictional incident timeline, redacted trace bundle, one disconfirmed false lead, three supported contributing conditions, containment, correction, evaluation/control/runbook changes, verification, and residual limitations.

Start from symptom, then branch by layer

Users experience symptoms: wrong candidate, unsupported certainty, duplicate effect, missing correction, unexplained abstention, or delay. Operators see symptoms too: behavior error, latency tail, queue, control rejection, or feedback category.

Monitoring practice distinguishes user-visible symptoms from internal causes. Cross-component trace semantics help correlate evidence. Neither creates causal certainty. [CLM-095]

Build a diagnostic tree across:

  • user task and interface;
  • product logic and metric;
  • data, source, label, and context;
  • learned component and version;
  • deterministic validator and policy;
  • tool and effect state;
  • orchestration and retry;
  • dependency and runtime;
  • permission, privacy, and security control;
  • instrumentation and sampling;
  • human review and organizational process.

For each hypothesis, record:

  • statement and layer;
  • predicted observations;
  • evidence for;
  • evidence against;
  • next discriminating test;
  • state: open, supported contributing condition, disconfirmed, or unresolved;
  • owner and time boundary.
Colorful three-dimensional layered diagnostic tree branching from user-visible symptom through product, data-context, model, control, tool, runtime, instrumentation, and human-process hypotheses, with short labels for evidence, disconfirmation, false lead, and next test.
F18.1 - Branch from the symptom and require layer-specific evidence before changing the system or naming a cause.

The tree protects against model-first diagnosis. It also protects against a new fashionable cause replacing careful investigation.

Contain consequence before optimizing explanation

Incident investigation and product optimization have different priorities. When consequential behavior continues, first reduce exposure and preserve state.

Containment options include:

  • disable one feature or configuration;
  • route a segment to control;
  • stop effects;
  • open a circuit;
  • quarantine data or a component version;
  • reduce capability;
  • abstain;
  • revoke permission;
  • increase qualified review;
  • restore a known bounded artifact.

Choose the narrowest control that stops consequence without creating equal or greater risk. Record owner and authority. Emergency action does not give the responder permanent risk authority.

Patchwork stops the bounded cohort at the first adjudicated critical compatibility pattern, restores deterministic search, and quarantines affected seller records. Optimization requests are rejected until containment is verified.

Preserve volatile evidence before it disappears, but do not copy sensitive content everywhere. Capture versions, trace IDs, evidence IDs, states, timestamps, control decisions, and hashes under the declared incident purpose.

Declare the incident boundary

An incident record needs:

  • detection source and declaration time;
  • affected behavior and user consequence;
  • population, segment, and time window;
  • current versions and exposure route;
  • known facts, unknowns, and assumptions;
  • immediate containment;
  • incident command and decision rights;
  • communication owners;
  • privacy and confidentiality constraints;
  • next evidence checkpoint.

Do not wait for a single root cause before declaring. Declaration creates coordination and authority. It should be reversible if evidence later narrows severity.

The Applied AI Engineer can supply system evidence and implement authorized containment. They do not automatically become incident commander, public communicator, privacy authority, or risk acceptor.

Build a chronological evidence timeline

Patchwork’s synthetic timeline is:

  1. 11:00 bounded cohort enabled;
  2. 11:08 engagement proxy rises;
  3. 11:12 first adjudicated compatibility report;
  4. 11:14 incident declared and cohort stopped;
  5. 11:18 seller-content quarantine enabled;
  6. 11:25 rollback replay passes frozen cases.

Each event links to an evidence identifier. The times are teaching fixtures, not real response metrics.

Separate event time, observation time, and record time. A report may describe an earlier interaction. A delayed queue can make detection appear later. Clock skew can invert events. Preserve uncertainty rather than forcing a clean story.

Include changes to telemetry and investigation itself. An analyst query, manual quarantine, or dashboard correction can alter what later observers see.

Timeline correctness matters because causal stories often rely on sequence. Metric rose before report does not mean the metric caused the incident. It identifies an order and next question.

Protect privacy during diagnosis

Incidents create pressure to log everything. That can amplify harm and still fail to explain behavior.

Patchwork’s trace allowlist keeps:

  • trace ID;
  • segment code;
  • behavior state;
  • failure layer;
  • evidence IDs;
  • policy flags;
  • latency.

It removes raw seller text and user query. Investigators first replay evidence identifiers against controlled synthetic or authorized sources.

Use staged access:

  1. aggregate symptom and segment;
  2. structured trace and versions;
  3. evidence identifiers and controlled state;
  4. synthetic replay;
  5. approved minimal content sample only if necessary;
  6. specialist or authority review.

Incident purpose does not erase retention, access, deletion, or disclosure obligations. Keep audit of break-glass access. Avoid copying raw content into chat, tickets, or public reports.

Test hypotheses through predictions and disconfirmation

A hypothesis earns attention by producing a discriminating prediction.

H1: model adapter regression

Prediction: paired replay should fail across records under the new model version and pass under the old.

Evidence against: model adapter is unchanged. Clean records pass. Failure appears only with corrupted seller data. H1 is a disconfirmed false lead.

H2: seller content escaped inert-data isolation

Prediction: failures cluster on records whose seller field enters features; removing or quarantining that field should restore the invariant.

Evidence for: SELLER-77 reproduces; the quarantine flag identifies the path; cleaned data passes replay. H2 is a supported contributing condition.

H3: engagement proxy rewarded unsupported attractiveness

Prediction: clicks rise while adjudicated compatibility errors also rise.

Evidence for: both occur in the synthetic fixture. The small cohort prevents a prevalence or causal estimate. H3 remains contributing, not universal.

H4: known-token control coverage was incomplete

Prediction: novel phrasing bypasses the pattern while other permission and tool controls hold.

Evidence for: the fixture bypasses CTL-01’s token pattern; permission and tool state remain intact. H4 is supported.

This process avoids both single-cause certainty and endless ambiguity. Multiple conditions can be supported at different strengths.

Look upstream before patching downstream

Data failures can cascade through organizational and technical stages. Weak definitions, provenance, ownership, and feedback can become model behavior symptoms later. [CLM-097]

In Patchwork, downstream ranking behavior is wrong because a seller field crossed an untrusted-data boundary. Tuning the ranker might suppress this record and leave the pathway open.

Ask:

  • Who created and owns the source?
  • Which transformations occurred?
  • Did provenance and trust level survive?
  • Which permissions and freshness state applied?
  • Did training, evaluation, and serving use equivalent definitions?
  • Did a cache or index preserve corrected state?
  • Which organizational incentive encouraged the field?
  • Did feedback route to the data owner?

Correction changes the upstream boundary and the downstream control, then replays affected cases.

Keep adversarial and misuse hypotheses in the tree

An incident may be accidental, adversarial, or both. Seller gaming can exploit ranking incentives without sophisticated model attack. Threat taxonomies help investigators consider poisoning, evasion, prompt injection, insecure output, privacy, and excessive agency. [CLM-098]

Do not label intent without evidence. Record seller manipulation hypothesis and observations. Preserve security-sensitive details in restricted evidence.

Also ask how the product rewarded the behavior. A seller can optimize for a metric because the system made the incentive valuable. Security and product design share the diagnosis.

Controls should not rely on detecting malicious intent. Enforce trust boundaries for all untrusted content.

Retries and partial failure can amplify incidents

During diagnosis, retries can increase load, duplicate observation, or repeat effects. Bounded retries, backoff, jitter, and idempotency remain active during incidents. [CLM-096]

Do not disable controls to get more data. If a dependency is impaired, open a circuit or reduce capability. If effect state is unknown, reconcile before retry. Record which layer owns attempts.

Investigation tools can also create load. A replay over a corrupted index may worsen saturation. Bound case count, concurrency, cost, and data scope.

If telemetry delivery fails, avoid synchronous logging retries on the user path. Expose evidence loss and decide whether affected behavior can continue.

Qualitative evidence can overturn aggregates

Aggregate engagement rose in Patchwork. Two adjudicated reports revealed a consequential category. This does not mean two reports estimate prevalence. It means the current aggregate was insufficient for the release claim.

First-party retrospectives show that positive user feedback and offline evaluation can miss a harmful behavior pattern and that rollback should be followed by evaluation and process change. [CLM-099]

Keep feedback selection visible. Investigate report mechanism, friction, and missing users. Add the discovered category to evaluation after provenance and contamination review.

Metric recovery is not user recovery. Click rate returning to baseline says nothing about people who received wrong compatibility guidance. Identify affected outputs through the smallest authorized evidence path and decide correction or communication with named owners.

Correct the system at multiple layers

Patchwork correction includes:

  • data: remove or quarantine instruction-like seller fields from ranking input;
  • control: replace token-only detection with structural allowed-field isolation plus adversarial fixtures;
  • code: enforce inert seller content before feature extraction;
  • evaluation: add a seller-manipulation release case;
  • observability: add a feedback-disagreement panel;
  • runbook: require a disconfirming model-version test;
  • ownership: route upstream data correction to the data owner.

A code patch alone is fragile. The same category can return through another component or later refactor.

Corrective action should change the system rather than blame an individual seller, reviewer, or engineer. Accountability still names owners and decision rights.

Verify recovery, not merely restoration

Restoration means the service or control route is available. Recovery means the affected behavior and state meet the bounded contract. Verification supplies evidence.

Patchwork verifies:

  • deterministic search restored;
  • affected seller records quarantined;
  • ordinary, long-tail, seller-manipulation, permission, and tool cases replayed;
  • CTL-01 implementation and monitor revised;
  • restricted segments remain restricted;
  • no effects occurred;
  • negative evidence and novel-manipulation residual remain;
  • release authority owns any renewed exposure.

One replay does not prove production prevalence or long-term recovery. Continue a bounded monitor window and preserve uncertainty.

Human-AI recovery includes expectation, correction, and behavior over time. Users may have changed trust, repeated work, or used another channel. [CLM-100]

Convert the incident into durable learning

An incident learning record includes:

  • what changed in the system;
  • what changed in evaluation;
  • what changed in control and monitoring;
  • what changed in runbook and ownership;
  • what remains unresolved;
  • which evidence verifies each change;
  • who owns follow-up;
  • when evidence expires;
  • conditions for renewed release.
Colorful three-dimensional incident learning loop from detection through declaration, containment, hypotheses, evidence and disconfirmation, correction, recovery, verification, and durable changes to evaluation, control, runbook, ownership, and release evidence, with short labels.
F18.2 - An incident closes only when verified learning changes durable artifacts and any renewed exposure returns to explicit authority.

Do not close because the metric is green or code is deployed. Close when containment, state reconciliation, affected-user decision, artifact changes, verification, owners, and residuals are recorded under the organization’s incident authority.

Write the incident report without false certainty

Use:

  • The incident involved interacting data-context, control, and product-metric conditions;
  • The model-version hypothesis was disconfirmed by unchanged version and paired replay;
  • Seller-content isolation failure is a supported contributing condition in covered evidence;
  • Novel manipulation remains possible.

Avoid:

  • The model caused the incident;
  • The root cause was one bad seller;
  • The issue is fixed forever;
  • No users were harmed without evidence;
  • publishing sensitive traces.

Complex systems often have contributing conditions rather than one root cause. Naming interaction improves correction.

Operate the incident as phases with explicit exits

Detect

Confirm that the signal exists and instrumentation is healthy. Record source, time, population, and uncertainty. Do not dismiss qualitative reports because aggregate metrics are green.

Exit: enough evidence to classify as expected behavior, investigation, or incident declaration.

Declare

Name incident scope, commander or authority, roles, communication channel, confidentiality, and immediate objectives.

Exit: decision rights and containment priority are explicit.

Contain

Stop or reduce consequence while preserving state. Choose rollback, kill, quarantine, abstention, circuit, or permission revocation.

Exit: affected behavior no longer expands within the declared boundary, or residual exposure is explicitly escalated.

Diagnose

Build timeline and layered hypotheses. Seek disconfirming evidence. Protect sensitive data.

Exit: supported contributing conditions are adequate to design correction, with unresolveds recorded.

Correct

Change data, code, configuration, control, process, or ownership. Do not optimize unrelated metrics.

Exit: corrections are implemented in a versioned candidate.

Recover

Restore acceptable behavior and reconcile state and user consequences.

Exit: the bounded behavior contract holds under replay and authorized observation.

Verify and learn

Test correction, preserve negative evidence, update artifacts, assign follow-up, and decide renewed exposure.

Exit: incident authority accepts closure criteria, not merely a green dashboard.

Phases can overlap. Containment can continue while diagnosis improves. The exits prevent a fast code deploy from skipping recovery.

Build an evidence bundle that survives handoff

The bundle should include:

  • incident and trace identifiers;
  • affected behavior and segment;
  • exposure and cohort route;
  • feature, config, model, data, schema, control, and index versions;
  • timeline with evidence links;
  • behavior, error, service, context, and cost signals;
  • relevant control decisions;
  • feedback provenance and selection limits;
  • hypotheses and predictions;
  • evidence for and against;
  • containment actions and authority;
  • correction commits or artifact identities;
  • replay and recovery results;
  • user repair decision;
  • residuals and follow-up.

Use a shared facts section. Team members can disagree on interpretation while preserving the same observed evidence. Mark inference, report, and confirmed state.

Do not attach broad raw-content exports. Use redacted structured traces and controlled links. Restrict security-sensitive reproduction details.

Hash or version fixtures so later replay is comparable. If incident data cannot be retained, create a permitted synthetic reproduction that preserves the mechanism and document the transformation.

Distinguish detection from declaration and severity

An alert detects a condition. A qualified role declares an incident under organizational policy. Severity reflects consequence, scope, reversibility, and operational needs. The engineer who writes the detector may not declare public severity.

Avoid severity based only on request count. One permission exposure or autonomous effect can be material. Conversely, high abstention can be correct containment.

Record why a severity changed. New evidence may narrow affected population or reveal delayed effects. Do not edit history; add the transition and authority.

Patchwork’s fictional exercise uses one bounded cohort and no external effects. It teaches process without claiming a real severity class.

Preserve authority during urgency

Urgency compresses time, not decision boundaries. Preassign:

  • incident command;
  • technical diagnosis lead;
  • release/rollback operator;
  • data and domain owner;
  • privacy/security reviewer;
  • user/support communication owner;
  • legal or regulatory escalation where applicable;
  • closure authority.

The Applied AI Engineer may recommend and execute a technical rollback under delegated authority. They should not improvise public disclosure, accept residual legal risk, or publish sensitive traces.

If the authorized person is unavailable, follow the predeclared escalation or safe default. Do not interpret silence as approval to continue exposure.

Ask one discriminating question at a time

During Patchwork diagnosis:

  1. Did the model adapter change? No.
  2. Does replay fail on clean and corrupted records? Only corrupted.
  3. Does removing seller text change the result? Yes.
  4. Did permission or tool controls fail? No.
  5. Did engagement and compatibility reports move together? Yes in the small fixture.
  6. Does structural allowed-field isolation stop reproduction? Yes in frozen cases.

Each answer changes hypothesis weight. This is stronger than a broad search through raw logs.

Prefer tests that separate alternatives. If both model and data changed, replay the same data across models and the same model across data versions. If instrumentation changed, compare independent evidence.

Document tests that cannot be run and why. Missing evidence is part of the incident conclusion.

Investigate metric gaming without moral shortcuts

Seller gaming can be intentional, emergent, or incentivized by product design. Examine:

  • which field influenced ranking;
  • whether sellers could observe the effect;
  • product messaging and incentives;
  • detection and enforcement consistency;
  • false positives for legitimate descriptions;
  • impact on shoppers and other sellers;
  • how engagement optimization rewarded the behavior.

Do not use bad actor as the whole cause. The system accepted the field, lacked structural isolation, and treated clicks as favorable.

Correction can include schema changes, field trust levels, review, incentive changes, control tests, and monitoring. Security action may be appropriate under policy, but it does not replace system correction.

Diagnose feedback disagreement

Create a view that compares:

  • engagement or preference;
  • behavior-contract states;
  • critical error cases;
  • evidence coverage;
  • correction and return reports;
  • affected segments;
  • versions and exposure;
  • report sampling and delay.

Do not blend into one score. The disagreement is the signal.

For Patchwork, rising clicks coexist with compatibility reports. One interpretation is attractive unsupported results. Alternative explanations include changed traffic, feedback prompt, instrumentation, or chance. The trace and replay support the data/control pathway in the fixture.

When a proxy conflicts with consequence, stop optimizing the proxy. Return to the user task and behavior contract.

Reconstruct affected state

Containment stops new consequence. The team must decide what already happened.

Use the smallest authorized evidence to identify:

  • exposed cohort and time window;
  • affected version and records;
  • outputs that crossed the interface;
  • corrections or user actions;
  • any external effect or unknown state;
  • support reports;
  • records requiring notification, reversal, or review.

Patchwork has no tool effects, but users may have seen wrong candidates. The fictional record says route adjudicated correction; it does not invent a production affected count.

If reconstruction is incomplete, state lower and upper bounds or unknown. Do not claim no impact because logs were minimized. Privacy-conscious observability should preserve decision-bearing IDs where justified, but blind spots remain.

Communicate without speculation

An internal update can contain:

  • confirmed facts;
  • affected scope currently known;
  • current containment;
  • hypotheses under test;
  • disconfirmed leads;
  • user and operational unknowns;
  • next evidence checkpoint;
  • decisions and authorities needed.

Avoid speculative root cause and exact resolution time. Explain what would change the estimate.

Separate technical detail by audience. Operators need action and state. Product/support need user behavior and repair. Security/privacy need restricted evidence. Executives need consequence, containment, uncertainty, and decision.

Keep shared facts consistent across versions. Do not soften a critical unknown in a high-level summary.

Correct data and provenance, not only features

For seller fields, define provenance and trust:

  • seller asserted;
  • platform verified;
  • derived by system;
  • domain reviewed;
  • unknown or conflicting.

The ranker may use seller-asserted descriptive text for relevance under safeguards, but compatibility claims require stronger evidence. Preserve trust through indexing and feature pipelines.

Quarantine changes need reindex and cache invalidation. Verify that old features do not survive. Update source documentation and owner. Add a monitor for trust-level loss.

If labels or examples were created from manipulated engagement, inspect training and evaluation contamination. Do not retrain blindly on the same proxy.

Improve the control from incident evidence

CTL-01’s first version combined known-token detection and structural projection, but the incident fixture assumes a pathway allowed seller text into features. The correction makes the structural boundary explicit before feature extraction and adds the novel phrase as an adversarial case.

Update the control chain:

  • threat path now includes ranking-feature injection;
  • implementation location moves earlier;
  • test covers novel phrasing and benign near-match;
  • monitor adds trust-level loss and quarantine rate;
  • owner includes data pipeline;
  • residual still says novel manipulation is possible;
  • release packet requires the revised version.

Do not claim the new pattern solves manipulation. The durable objective remains inert untrusted content.

Verify with positive, negative, and recurrence tests

Positive: legitimate seller records still support allowed discovery.

Negative: instruction-like text cannot influence policy, compatibility, or tool state.

Boundary: benign uses of words like ignore do not automatically remove required structured evidence.

Recurrence: the exact incident fixture fails before correction and passes after.

Interaction: permission, stale evidence, fallback, audit, and rollback still work.

Load: quarantine and reindex do not exceed resource envelope.

Recovery: control route handles traffic and affected traces reconcile.

Preserve before/after versions and negative evidence. If a test is changed to match implementation, explain why the contract changed.

Include user recovery in closure

Human-AI interaction can change expectations and trust. Ask:

  • Could users detect the wrong behavior?
  • Was correction available and understandable?
  • Did they repeat or abandon the task?
  • Could they have acted on the result elsewhere?
  • What notification or repair is authorized?
  • Does the interface need a changed limitation or evidence display?

Do not use restored engagement as evidence of recovered trust. Qualitative follow-up or longitudinal behavior may be needed under privacy constraints.

The incident authority decides the repair plan with product, support, domain, privacy, and legal roles where applicable.

Write corrective actions that change conditions

Weak: Train engineers to be careful with seller text.

Stronger: Schema projection excludes untrusted seller free text before feature extraction; integration test PW-ADV-17 asserts no policy or compatibility influence; data owner monitors trust-level loss.

Every action names:

  • condition changed;
  • owner;
  • artifact or system location;
  • due evidence event;
  • verification;
  • residual;
  • dependency;
  • closure authority.

Training can support a control, but it should not be the only correction for an enforceable system boundary.

Prioritize actions by recurrence and consequence, not visibility. Avoid a long list that no one owns.

Hold an evidence-centered review

Begin with user consequence and timeline, not a polished root-cause story. Walk the architecture and hypotheses. Invite data, product, domain, operations, security, privacy, and support perspectives.

Use blameless analysis to examine system conditions while still assigning accountable owners. Avoid speculation about individuals’ motives.

Review what went well, what limited consequence, what delayed detection, and where authority was unclear. Preserve dissent. A reviewer may believe the product metric remains invalid even after technical repair; record and route that decision.

End with artifact changes, owners, verification, residuals, and renewed-release conditions.

Work the false lead deliberately

Model blame is attractive because the system contains a learned component and the visible result is wrong. Test it rigorously.

Freeze the failing request and evidence versions. Run the same corrupted record through current and previous model adapters. If both fail, adapter change is less likely. Run clean and corrupted data through the same adapter. If only corrupted fails, data/context becomes more plausible.

Inspect deterministic validators. Did the output cross because a control was absent, misconfigured, or bypassed? Replay with structural isolation enabled. If it passes, the boundary is discriminating evidence.

Check instrumentation. A version field can be stale. Verify artifact identity from the release record, not only a dashboard label.

Disconfirming H1 does not prove H2 exclusively. Product metric and control coverage remain contributing conditions. Record confidence and residual alternatives.

This explicit false-lead exercise makes intellectual humility executable.

Test the instrumentation hypothesis

Could compatibility reports have risen because the feedback prompt moved? Could engagement have risen because click instrumentation duplicated? Could segment labels have changed?

Compare event schema and producer versions. Check expected request-to-event ratios, duplicate IDs, sampling, queue delay, and independent support records. Replay dashboard query on frozen events.

If instrumentation error exists, correct it and reinterpret the window. Do not discard user reports merely because one metric is wrong.

Patchwork’s fixture assumes instrumentation holds. A real incident record must show the check.

Test the distribution-change hypothesis

The cohort may contain more long-tail seller records than offline evaluation. Compare covered distributions by source type, trust level, revision, and seller-field presence.

Distribution difference can explain exposure without being the initiating defect. The control should still prevent untrusted instruction influence. Record both: real-use mix revealed the path, and structural isolation failed.

Do not label every distribution change drift or cause. Ask whether behavior changes under paired controlled data.

Add the newly observed slice to evaluation without claiming prevalence from the incident sample.

Handle an incident with an unknown effect

Patchwork’s main incident has no effects, but the runbook must cover them.

If a seller message submission times out:

  1. stop new submissions for the affected operation;
  2. preserve operation key and confirmed payload hash;
  3. query authoritative effect ledger;
  4. classify completed, known-not-started, failed-known, or unknown;
  5. never retry unknown;
  6. reconcile UI and user expectation;
  7. identify duplicate or missing effects;
  8. apply authorized repair;
  9. add regression and runbook evidence.

Retries and incident investigation can otherwise amplify consequence. Keep the effect owner and authority explicit.

Protect the incident channel

Set rules before crisis:

  • no raw prompts, images, seller text, secrets, or direct identity in general channels;
  • use opaque incident and trace IDs;
  • restricted links expire;
  • exports require purpose and approval;
  • screenshots follow the same classification;
  • external communication has named authority;
  • retention and deletion cover chat and tickets;
  • adversarial content remains inert in investigation tools.

A prompt injection string copied into an AI-assisted incident tool can create a second attack path. Treat trace strings as data, escape them, and use structured summaries.

If a public report is required, separate transparency from exposing sensitive reproduction details. That decision is outside the engineer’s unilateral authority.

Diagnose reviewer and organizational conditions

Why did the first report reach adjudication? Why did the second strengthen the pattern? Did reviewers have evidence and time? Were earlier signals ignored because engagement was positive?

Possible conditions:

  • no owner for disagreement panel;
  • support categories too broad;
  • reviewer queue delayed;
  • launch incentive favored clicks;
  • data owner absent from release review;
  • model-first culture narrowed diagnosis;
  • residual was documented but not propagated.

These are system conditions, not excuses. Correct ownership, routing, and decision process where evidence supports them.

Separate correction, recovery, and prevention

Correction removes the immediate defective condition: quarantine seller field and enforce structural projection.

Recovery restores bounded acceptable behavior: control route active, affected state reconciled, cases pass, user repair decided.

Prevention reduces recurrence: control chain revised, data trust preserved, evaluation case added, release gate and monitor changed, ownership clarified.

One action can contribute to several, but the evidence differs. A code merge is correction evidence, not user recovery.

Define verification before seeing the fix

Predeclare:

  • incident fixture must fail on old version and pass on candidate;
  • ordinary discovery quality cannot regress below floor;
  • permission and tool controls remain zero-tolerance;
  • benign seller descriptions remain usable;
  • quarantine does not create unbounded tail or queue;
  • reindex removes old feature state;
  • redacted telemetry shows control decision;
  • restricted segments stay restricted;
  • affected-output review is complete within declared evidence.

If verification criteria change, version and justify them. Avoid making the test easier after a preferred fix fails.

Decide renewed exposure through PF-11 v0.1

Incident closure does not automatically relaunch. Return to readiness:

  • update feature, data, control, evaluation, and runbook versions;
  • attach negative incident evidence;
  • show correction and replay;
  • reassess residual manipulation risk;
  • confirm privacy and security review;
  • rehearse stop and rollback;
  • choose scope and cohort anew;
  • obtain authority.

The first renewed stage may be shadow, not the previous cohort percentage. The next uncertainty is whether structural isolation holds under live-shaped seller distributions, not whether engagement recovers.

Build a durable learning ledger

Each item contains:

  • observation;
  • evidence and confidence;
  • changed assumption;
  • artifact affected;
  • owner;
  • verification;
  • residual;
  • reuse boundary;
  • review date.

Patchwork entries include:

  • seller free text can enter ranking features unless structurally excluded;
  • engagement can disagree with compatibility consequence;
  • unchanged model version is useful disconfirmation;
  • feedback needs a critical path to release stop;
  • data owner belongs in correction;
  • runbook needs a model-false-lead step.

Do not generalize one fictional incident to all marketplaces. The ledger preserves local mechanism and bounded reuse.

Verify learning after time passes

At the next relevant change or exercise, ask:

  • Does the incident fixture still run?
  • Is the control invoked on every ingestion path?
  • Does the monitor fire?
  • Can the on-call find the runbook?
  • Is the data owner still assigned?
  • Does release review see the residual?
  • Are privacy controls and deletion working?
  • Did a new provider or schema invalidate evidence?

An action marked complete but never exercised is weak evidence. Rehearsal turns documentation into operational memory.

Measure incident response without gaming it

Time to detect, contain, recover, and verify can be useful, but definitions matter. Optimizing declaration time can discourage careful classification; optimizing closure time can hide residual work.

Report phases with boundaries and uncertainty. Include user consequence, recurrence, evidence quality, and action completion. Do not reward fewer incidents if reporting becomes harder.

The fictional Patchwork timestamps show sequence only. They do not set an objective.

Common incident disputes

“The model is probabilistic, so this is expected.” Expected uncertainty still needs bounded behavior and controls. Diagnose the failed contract.

“Clicks increased, so users preferred it.” Clicks are a proxy. Compatibility reports reveal conflicting consequence.

“One bad record caused it.” The record activated a path the system permitted. Correct the boundary and incentives.

“We rolled back, so the incident is over.” Reconcile affected state, verify behavior, decide user repair, and change artifacts.

“We need raw logs to know.” Start with structured traces, versions, evidence IDs, and replay. Escalate minimal content under authority only if necessary.

“The fix passed, so relaunch.” Return to readiness and authority with updated evidence.

“Root cause must be singular.” Record supported interacting conditions and disconfirmed alternatives.

Exercise the layered bundle

Give responders:

  • green service dashboard;
  • rising engagement;
  • two adjudicated compatibility reports;
  • unchanged model version;
  • corrupted seller record ID;
  • redacted trace flags;
  • permission and tool controls holding;
  • old and new data replay;
  • one failed known-token detector;
  • cohort and rollback packet.

Ask them to:

  1. declare facts and unknowns;
  2. contain before analysis;
  3. write at least four layer hypotheses;
  4. identify the model false lead;
  5. propose discriminating tests;
  6. protect raw content;
  7. correct data, control, and code;
  8. add evaluation, monitor, runbook, and ownership changes;
  9. verify recovery;
  10. route renewed release to authority.

Pass only when the response does more than patch ranking.

Create an incident decision log

Record each material decision with:

  • timestamp and decision maker;
  • authority or delegation;
  • facts available;
  • unknowns and hypotheses;
  • options considered;
  • chosen action and expected consequence;
  • rollback or reversal;
  • next review trigger.

Examples:

11:14: stop cohort because adjudicated compatibility reports meet the critical behavior trigger; authority is the named incident/release role; control route is available.

11:18: quarantine seller records because paired evidence localizes failure to seller-content feature path; data owner will review broader affected set.

11:25: do not relaunch despite replay passing because control and release packets require versioned review and authority.

The log prevents hindsight from making every choice look obvious. It also shows when authority or evidence was missing.

Handle evidence conflicts

Suppose service traces show no error, model evaluation passes, and users report wrong compatibility. Do not choose one source by prestige. Map each claim.

Service traces support availability and latency. Model evaluation supports covered frozen cases. User reports support those selected interactions. Domain review supports compatibility criteria. Data replay supports the seller-path mechanism.

The conflict dissolves when evidence is assigned to its actual claim. The system can be available, pass old evaluation, and still fail a new critical behavior category.

When sources directly disagree on the same fact, inspect identity, timestamp, transformation, and authority. Preserve unresolved status if evidence cannot reconcile.

Define incident closure criteria

Require:

  • exposure contained;
  • state and effects reconciled;
  • affected population bounded as far as evidence permits;
  • user repair decision assigned;
  • supported contributing conditions and disconfirmed leads recorded;
  • corrections implemented and versioned;
  • regression, control, and recovery verification passed;
  • privacy and security handling complete;
  • evaluation, monitor, runbook, and ownership updated;
  • residuals and follow-up owners recorded;
  • renewed release separated from closure;
  • closure authority disposition.

If one item remains open, the incident can move to follow-up state without pretending it is complete. Operational severity may reduce while corrective work continues.

Preserve selected reports responsibly

A report that becomes an evaluation case needs provenance, permission, minimization, transformation, split assignment, and retention. Remove identity and content not needed for the behavior mechanism. Consider whether the reporter expected this use.

Keep incident-derived cases away from development paths that could contaminate release evaluation. Version the suite and document why the case was added.

Do not retain every report forever. Aggregate categories when detailed purpose ends, or delete under policy. A synthetic reconstruction can preserve mechanism where raw content is unnecessary.

Feedback selection remains a limitation. The case proves possibility in scope, not frequency.

Examine near misses

A control rejection that prevents user consequence can still reveal a dangerous path. Record near misses when:

  • permission bypass reached a late validator;
  • unsupported tool proposal reached confirmation preview;
  • seller injection changed ranking but output was caught;
  • duplicate effect was reconciled before user harm;
  • restricted segment almost entered cohort.

Near misses can update architecture and controls without waiting for incident. Avoid inflating counts as harm metrics; keep state and consequence distinct.

Review whether detection was early enough and whether the same path could bypass another boundary.

Include recovery load and secondary failure

Rollback, quarantine, reindex, replay, and affected-state search consume capacity. Budget them. An aggressive reindex can overload the control path. A broad query can expose sensitive content. A cache flush can create tail spikes.

Stage recovery:

  • protect current user traffic;
  • rate-limit background repair;
  • use permission-aware queries;
  • monitor queues and tails;
  • verify partial completion;
  • keep circuits and retries bounded;
  • stop if recovery creates new consequence.

Test a secondary failure during rehearsal. If quarantine service times out, default to excluding the affected field or record, not restoring it.

Final incident questions

  • What user-visible consequence triggered response?
  • Is instrumentation healthy?
  • What is currently contained and what still expands?
  • Which authority owns stop, communication, repair, and closure?
  • Which layers have open hypotheses?
  • What observation would disconfirm each?
  • Is model blame supported or merely salient?
  • Are adversarial and product-incentive paths considered?
  • Are traces minimal and access controlled?
  • Did retries or recovery amplify load or effects?
  • What changed beyond code?
  • How is user recovery evaluated?
  • Which residual blocks renewed exposure?
  • Can a later team replay the learning?

The quality of an incident response appears in these distinctions, not in the confidence of a root-cause sentence.

Also ask what would have detected the condition earlier with less user exposure. The answer might be a seller-trust evaluation slice, a structural ingestion invariant, a disagreement panel, a support escalation rule, or a shadow comparison. Add the smallest useful signal or case. Do not respond by collecting all content.

Finally, verify that ownership survives handoff. The data owner accepts the source correction, the control owner maintains isolation, the evaluation owner keeps the new case, the runbook owner rehearses the false-lead step, and release authority sees the residual before any cohort returns. A list of actions without durable owners will decay into the same vulnerability.

Incident learning should reduce future uncertainty, not merely create a long document. Link every conclusion to an artifact, test, owner, and review trigger. If a conclusion cannot change the system or a decision, explain why it belongs in the record.

Preserve the distinction between not observed, ruled out in covered evidence, and impossible. The model hypothesis is disconfirmed for this fixture because version and paired replay contradict it. That does not prove models can never contribute to similar incidents. Likewise, permission and tool controls holding in the observed traces does not establish their universal effectiveness.

Use closure language that a later investigator can safely reuse: scope, versions, evidence, confidence, residual, and trigger. Avoid dramatic certainty that will be copied into future decisions without its limitations.

Archive the redacted bundle where authorized teammates can find it, not in one person’s local notes. Include reproduction commands for synthetic fixtures, expected outputs, and the reason each artifact changed. Test those commands after the incident branch merges.

When the next incident looks similar, begin with the old hypotheses but do not assume identity. Compare versions, population, symptom, and evidence. Reuse method and fixtures carefully; reopen cause and consequence.

Preserve those distinctions.

What the companion proves

Run:

node --test content/publications/applied-ai-engineering/companion/tests/chapter-18.test.mjs

The tests establish that PF-11 v0.2 has a chronological timeline, a disconfirmed false lead, three contributing conditions, privacy-redacted traces, containment before optimization, and durable evaluation/control/runbook changes.

They do not establish a real incident, real affected population, causal certainty, legal response, complete threat coverage, user recovery, or production authorization.

Carry incident discipline into planned change

Real use teaches by exposing where evidence and assumptions break. The learning is durable only when it changes system boundaries, evaluation, controls, monitoring, ownership, and future release decisions.

Patchwork’s incident demonstrates the method: begin with user-visible consequence, stop exposure, preserve minimal evidence, test competing layer hypotheses, disconfirm the attractive model story, correct upstream and downstream conditions, replay, and keep residuals.

Chapter 19 applies the same discipline to a planned or forced provider change. Migration is not merely swapping an API. It is a new behavior, control, cost, observability, and recovery claim.