Stabilize Under Real Conditions
Operate inside customer incident command, contain impact, communicate calibrated facts, restore service, correct systemic conditions, and prove stabilization exit.
When reality contradicts readiness, first protect people and the workflow. Then establish truth.
The FDE may be the deepest technical lead, but operates inside the customer’s incident structure. Incident command, customer/public communications, legal/privacy/security/safety decisions, and business authority remain explicitly designated. [CLM-142] [CLM-151]
Establish command and impact before deep debugging
Confirm:
- incident commander and deputies;
- technical, customer-impact, and communications leads;
- affected users/cohorts/workflow/data and consequence;
- current exposure and containment authority;
- evidence/timeline/decision log location;
- update cadence, escalation, and next checkpoint;
- access, privacy, safety, legal, and support constraints.
Do not begin with “root cause.” Initial priorities are impact, command, containment options, and evidence-safe coordination.
The companion rejects a record where an FDE silently names themselves commander without explicit designation.
Write the first ten-minute record
Before a deep technical theory, record who declared the incident and who commands, the affected workflow/cohorts, observed and worst credible near-term consequence, current exposure, containment options, first evidence requests, next update, and authorized communicator.
For a fictional Orchid low-connectivity cohort:
At 10:07 support confirmed technicians at two sites cannot complete evidence review within the operating window. No prohibited approval is confirmed. Expansion is held. Customer operations commands; the FDE leads technical diagnosis. Current unknowns are site connectivity, retry amplification, stale configuration, and evidence-service latency. Offline fallback remains available but adds manual time. Next update is 10:25.
This is safer than system slow, engineering investigating. It states impact, boundary, containment, uncertainty, and cadence without inventing cause.
Run parallel incident tracks
Impact and command
Who is affected, what can/cannot be done, how consequence changes, which authority decides.
Diagnosis and evidence
Correlated metrics, traces, events, workflow/data/model/policy/support evidence; hypotheses and falsification.
Containment and recovery
Disable/hold/isolate/fallback/reconcile/rollback/roll-forward/restore/stop, then verify user-visible state.
Communications and support
Confirmed impact/actions/unknowns/next cadence through authorized owners; support scripts and escalation.
Correction and learning
Systemic condition, durable change, regression evidence, monitoring window, owners/dates.
These parallel tracks are original synthesis, not a universal incident standard. [CLM-141]
Coordinate tracks through decisions
Parallel does not mean independent. Each track reports decisions and evidence through command.
If diagnosis finds retry amplification, containment must choose whether to isolate the cohort, disable suggestion, cap retry work, or activate fallback. Communications explains user effect without exposing speculation. Evidence preservation captures version/correlation state before configuration changes. Correction records the lasting system change after recovery.
Use a decision log entry with timestamp, decision, facts/unknowns, options, consequence, authority, action owner, review trigger, and linked evidence. The log prevents a fast response from becoming an unauditable sequence of chat messages.
Keep a confidence-tagged timeline
Each entry includes time/timezone, track, statement, confidence, evidence, actor/decision, and correction history.
- confirmed: supported by named evidence;
- probable: current hypothesis with basis and test;
- unknown: important unresolved question.
Do not rewrite old entries to look certain. Append corrections and decision changes. The companion requires evidence for a confirmed entry and preserves chronology.
An example sequence:
10:07 confirmed: two sites exceed the evidence-review latency window; linked support cases.10:10 confirmed: expansion held by release owner.10:13 probable: retry demand may amplify evidence-service saturation; trace sample supports correlation, capacity evidence incomplete.10:17 decision: isolate two sites and cap new suggestion work; operations authority; review at 10:30.10:23 disproved: stale configuration not present; release/config hashes match.10:28 confirmed: retry queue drains after isolation; user fallback remains available.
The timeline preserves the disproved hypothesis because it explains why the team stopped pursuing it.
Metrics, traces, logs, workflow state, data/model/policy, and support answer different questions. [CLM-145]
Communicate calibrated facts
An update states:
- observed impact/scope;
- confirmed facts and evidence time;
- unknowns/hypotheses labeled;
- containment/restoration actions and consequence;
- next update time;
- approved workaround/support path;
- ETA only when authorized and genuinely supported.
The companion rejects a “probable” ETA presented as ETA. [CLM-143]
Use a fixed update structure
Write impact, confirmed facts, labeled unknowns/hypotheses, actions/owners, user fallback, next cadence, and authority in that order. If no credible restoration estimate exists, say so and give the next evidence checkpoint. A confident cadence is not a speculative resolution time. Customer/public communications remain with the authorized owner. [CLM-151]
Avoid speculative cause, blame, invented reassurance, and sensitive technical detail in broad channels. Communications use the customer-approved owner/process.
Diagnose by consequence and boundary
For Orchid low-connectivity review failure:
- confirm affected cohort and whether safety workflow is blocked/misleading;
- hold expansion and enable truthful fallback;
- compare release/config/site/dependency and retry/new-demand signals;
- trace equipment/evidence/inventory/model/policy/approval path;
- inspect capacity, stale configuration, timeout and retry amplification;
- preserve minimized evidence and reconcile unknown intents;
- choose containment/recovery using Chapter 14’s state tree.
For an AI-quality incident, service health can remain green. Investigate equipment binding, evidence freshness, retrieval, model/prompt/index, policy, reviewer workload, critical segments, and support reports. [CLM-150]
Follow a boundary-first diagnostic sequence
Ask which boundary can produce the observed state: user/device/connectivity; identity/tenant/region; interface semantics; dependency capacity/finality; retrieval/provenance; model/prompt/policy/reviewer; release/config/migration; or telemetry/support measurement.
For each candidate, name the next discriminating evidence before broad searching. Compare affected and unaffected cohorts, versions, and workflow states. Preserve tails and critical segments. Metrics show magnitude/trend; traces show path/timing; logs show selected events; workflow state shows user consequence; support reports show lived symptoms. None is a universal truth source. [CLM-145]
Balance evidence preservation and containment
Preserve release/config/policy/model identifiers, correlation/intent IDs, timeline, decisions, minimized telemetry, relevant artifacts, and access/audit state.
Do not delay necessary harm reduction merely to collect perfect evidence. Do not casually copy sensitive payloads or alter/delete evidence. Follow customer incident, privacy, security, safety, and legal procedures. [CLM-146]
Use an evidence-preservation decision: what is needed to understand or prove the state, what is safe/legal to retain, who can access it, and whether collection delays containment. Preserve identifiers, versions, decisions, and minimized diagnostic state before copying raw content. If an authorized forensic process exists, follow it; the FDE does not improvise specialist collection.
When immediate shutdown would erase volatile evidence but continued operation can harm users, consequence comes first. Record the lost evidence and why containment was chosen. Honest uncertainty is preferable to additional harm for a perfect timeline.
Separate restore from correct
Restore: return an acceptable user-visible state: disable feature, isolate cohort, fallback, rollback/roll-forward, restore, reconcile.
Correct: remove/mitigate contributing system conditions, verify regressions, update controls/signals/runbooks/ownership, and monitor.
A workaround can restore quickly and remain dangerous as permanent architecture. Restoration does not end stabilization. [CLM-144]
For the Orchid example, isolating two sites and activating manual evidence review may restore an acceptable user path. It does not correct retry amplification, weak capacity gating, or the signal that failed to distinguish new demand from retries. The restoration record states current limitations and support burden. The corrective record changes budgets, queue behavior, tests, alerts, and the runbook, then verifies the affected cohort.
Define restoration acceptance separately from stabilization exit. Restoration may require an acceptable bounded workflow now; stabilization requires durable correction and evidence over time.
Find contributing conditions, not a person to blame
Ask how design, review, data, defaults, access, alerts, rollout, workload, incentives, documentation, ownership, or recovery made the event possible or harder.
Blameless learning is not accountability-free. It preserves evidence, actions, owners, local legal/safety/employment obligations, and real decision responsibility. [CLM-147]
Avoid stopping at a label such as operator error, bad deploy, model hallucination, or customer network. Ask what system conditions made the action possible, hard to detect, or hard to recover: confusing defaults, missing review, ambiguous ownership, inadequate capacity test, hidden configuration, poor evidence display, unavailable access, incentive conflict, or weak rollout.
A contributing-condition analysis can include workflow, software, data, environment, controls, tools, documentation, staffing, communication, and decision rights. It should not speculate about individuals. Where misconduct, safety, legal, or employment processes apply, route them to the designated authority rather than using a technical post-incident review as a substitute.
Write corrective actions as system changes
Each action contains:
- contributing condition/consequence;
- system/control/test/signal/runbook/process change;
- owner/due date/priority;
- verification method/data/environment;
- observed evidence and residual gap;
- escalation if overdue.
“Be more careful” is not a corrective system action. [CLM-149]
For retry amplification: separate new/retry budgets, preserve intent/reconciliation, cap/queue/fallback, add cohort capacity tests/signals, and verify failure recovery.
For stale configuration: version/promote/validate config, fail closed, correlate release state, add drift signal and rollback/roll-forward exercise.
For quality regression: restrict model/cohort, restore known bundle/fallback, add critical cases/segments/graders, verify policy/review, monitor the correction.
Verify corrective action, not ticket closure
For retry amplification, a completed action is not added limit. It includes a stable intent/reconciliation contract, separate new/retry capacity, bounded queue/fallback, a representative failure test, observable cohort signals, an owner, and evidence that recovery no longer amplifies load under the tested condition.
For stale configuration, verify version binding, promotion checks, drift detection, fail-closed behavior, and recovery with the affected environment. For a quality regression, bind model/prompt/index/policy/grader versions and confirm critical segments plus human-review behavior. Close only when the observed evidence meets the stated pass condition.
Define stabilization exit evidence
Require:
- impact contained and affected state reconciled;
- service/workflow restored under current scope;
- durable correction implemented and reviewed;
- regression/failure/recovery evidence passes;
- monitoring window complete across affected cohorts/tails;
- support/adoption state understood;
- corrective actions ownered/tracked;
- customer operators own current runbooks/access;
- residual risk/exception formally disposed;
- communications closeout authorized.
The companion refuses exit when only impactContained and serviceRestored are true. [CLM-148]
Conduct a timed stabilization exercise
Inject a low-connectivity latency increase followed by retry amplification. Give participants changing evidence: first a support report, then healthy aggregate availability, then cohort latency, then a trace sample, then evidence that configuration is current.
Evaluate whether the team:
- establishes customer command and technical roles;
- states impact and holds expansion;
- keeps confirmed/probable/unknown distinct;
- chooses next discriminating evidence;
- contains consequence without destroying evidence casually;
- communicates fallback and cadence without speculative ETA;
- restores an acceptable workflow;
- finds systemic contributing conditions;
- writes verifiable corrective actions;
- refuses stabilization exit until correction, regression evidence, observation, ownership, and residual-risk disposition exist.
The output is the timeline, decision log, updates, containment/recovery record, contributing conditions, corrective actions, verification window, and stabilization decision. Speed matters, but coherence and safe judgment matter more.
Failure modes and repairs
Debugging starts before command
Repair: name impact, commander, leads, exposure, containment authority, log, and cadence first.
A hypothesis becomes external fact
Repair: preserve confidence labels, evidence, corrections, and authorized communications. Do not promise an unsupported ETA. [CLM-143]
Healthy infrastructure hides workflow harm
Repair: inspect workflow, data, model/policy, approval, cohort, and support evidence as well as service metrics. [CLM-145] [CLM-150]
Restoration closes the incident
Repair: keep stabilization open through durable correction, regression evidence, observation, ownership, and residual-risk disposition. [CLM-148]
Blame replaces explanation
Repair: identify systemic contributing conditions and evidence-bearing corrective actions while preserving real accountability and formal processes. [CLM-147]
A permanent workaround becomes the architecture
Repair: record its limitations/expiry and deliver the durable system change or an explicit accepted exception.
Assemble the stabilization report
The report is not a polished story that removes uncertainty. It is the versioned bridge from incident state to owned operation.
Include:
- incident ID, scope, cohort, duration, and user/workflow consequence;
- designated command and communications authority;
- confidence-tagged timeline and decision log;
- releases/config/model/policy/dependency states involved;
- containment and restored user path;
- reconciled external/data state and remaining unknowns;
- contributing conditions across workflow, system, controls, and ownership;
- corrective actions with verification evidence;
- observation window across affected segments and tails;
- support/adoption impact and communication closeout;
- residual risk/exception and formal authority;
- operator/runbook/access transfer;
- stabilization decision and review trigger.
Continue the Orchid scenario through exit
At 10:17 the two affected sites are isolated and new suggestion work is capped. At 10:28 the retry queue drains; manual fallback works but adds delay. By 10:50 the team confirms that clients treated unknown inventory completion as new requests after a connectivity timeout. No wrong evidence binding is observed. The restored path is acceptable for the bounded cohort, but expansion remains held.
Correction separates new and retry capacity, preserves intent IDs, forces reconciliation before another external effect, adds queue/fallback limits, and instruments low-connectivity plus retry demand. A regression suite injects lost responses and verifies one external effect. A failure rehearsal verifies recovery and a customer operator executes the updated runbook.
The observation window includes both affected sites, sufficient unknown-completion injections, support load, latency tails, and reconciliation age. One known limitation remains: the real inventory provider’s rare finality behavior is not represented by the local double. The integration owner schedules representative evidence before wider exposure.
Stabilization can close for the current bounded path if the designated authority accepts that limitation and the existing scope does not depend on the unsupported case. It cannot close for a wider claim. This distinction preserves value without laundering a gap.
Transfer the method to a non-AI incident
Consider a fictional public-sector document migration where a partial schema rollout makes some records unreadable. The same method applies: establish command/impact, expose divergent state, stop mutation, select recovery by compatibility and consequence, communicate verified user paths, correct the migration/control, and prove restored records plus operator ownership.
The diagnosis does not mention models, prompts, or graders because they are irrelevant. What transfers is confidence-tagged evidence, actual-state recovery, designated authority, systemic correction, and stabilization exit. This case checks that the chapter teaches deployment operations rather than an AI-specific incident playbook.
Review the timed exercise
After the simulation, ask participants to identify:
- which early fact changed containment;
- which hypothesis was disproved and how;
- whether any update overstated certainty or authority;
- what evidence was lost or protected during containment;
- where restoration was mistaken for correction;
- whether corrective actions have verifiable system effects;
- which criterion actually permitted stabilization exit;
- what remains unknown and who owns it.
A fast containment with an incoherent log needs repair. A precise log that delays harm reduction also fails. Professional performance balances consequence, evidence, authority, and time.
Complete OA-10 stabilization layer
The artifact contains command/roles, impact, chronological confidence-tagged timeline, decision log, customer updates, evidence, containment/restoration, contributing conditions, corrective actions, verification window, residual risk, and stabilization disposition.
Incident response connects detection/response/recovery/improvement to organizational risk management. [CLM-152]
The companion adds five tests for evidence-backed chronology, command boundary, speculative ETA rejection, systemic corrective action, and restoration-versus-stabilization.
The Chapter 16 gate
Stabilization exits only when:
- customer command/communications authority is explicit;
- impact/facts/unknowns/decisions/timeline remain auditable;
- containment and recovered user state are verified;
- evidence is preserved without delaying harm reduction;
- durable correction and regression evidence exist;
- observation covers affected cohorts/quality;
- corrective actions and residual risks are ownered;
- customer operators/support accept the current state;
- companion incident tests pass.
Chapter 17, Make the Solution Usable and Ownable, tests whether users can succeed and customer operators can release, recover, support, and decide without hidden FDE dependency.