NEWProduction Web Themes & Turnkey ArchitecturesGet Lifetime Pass ($199) →
KNKomal Nakrani
Get All Access
ThemesDocsAll-Access PassGet All Access ($199)
Book overview
15/Forward Deployed Engineering

Cross the Production Threshold

Use readiness evidence, unresolved risk, authority, cohort design, signals, communications, and recovery triggers to choose go, conditional go, reduced scope, delay, or stop.

Production is not a place. It is exposure to real users, data, dependencies, authority, support, and consequence.

The readiness decision is therefore not “Is the build green?” It is:

Does the current evidence justify this exact cohort and duration under these conditions, signals, owners, communications, and recovery triggers?

The answer can be go, conditional go, reduced scope, delay, or stop. A review that permits only “go” is ceremony. [CLM-134]

Assemble the readiness evidence packet

Bring forward the authoritative artifacts:

  • outcome/guardrail/decision-rights contract (OA-03);
  • scope, dependencies, uncertainty, blast radius (OA-04);
  • architecture/interfaces/environment/AI/governance (OA-05);
  • runnable slice and debt (OA-06);
  • verification/evaluation/UAT/limitations (OA-07);
  • signals/objectives/alerts/runbooks/capacity/cost/release/recovery (OA-08);
  • open risks, exceptions, owners, authority, expiry;
  • user/operator/support/accessibility preparation;
  • cohort/ramp/observation/communication/command schedule.
Readiness gate where outcome, scope, design, implementation, verification, operations, recovery, ownership, and risk evidence receive met, gap, exception, reduced-scope, or block dispositions.
F15.1 - Readiness is an evidence decision. Named authority chooses go, conditional go, reduced scope, delay, or stop from explicit dispositions.

For each criterion record status:

  • met: current executed evidence supports the criterion;
  • gap: evidence missing, stale, inconclusive, or failed;
  • exception: designated owner accepts stated residual condition under scope/conditions/expiry;
  • reduced scope: evidence supports only a narrower cohort/capability;
  • block: consequence or authority rule prohibits exposure.

Missing evidence is never silently “not applicable” or passed. [CLM-130]

The five-way disposition and packet are original synthesis tools. [CLM-129]

Build a criterion table that can reject the launch

Every row needs criterion, consequence, required evidence/version, observed evidence, freshness/coverage, disposition, owner/authority, condition, and next trigger.

For fictional Orchid:

Criterion Evidence state Disposition
equipment/evidence binding across clear, ambiguous, and missing states current contract and slice tests pass met for bounded pump family
prohibited suggestion blocked before approval critical-segment evaluation passes; limited production-like volume met for internal cohort, uncertainty recorded
qualified review representative internal task evidence passes met for internal cohort only
low-connectivity fallback synthetic test passes; target field-site evidence missing reduced scope; exclude low-connectivity sites
break-glass recovery under failed identity dependency not executed gap; blocks safety-relevant/wider cohort
support/command schedule named and rehearsed for 48-hour window met under stated window

The table prevents a green aggregate from hiding an unexecuted consequential control. It also lets supported value continue in a narrower cohort without pretending the gap disappeared.

Separate facts, unknowns, judgments, and preferences

Facts

Versioned observations: tests passed/failed, restore executed/not executed, cohort signal coverage, known latency/capacity, approvals, artifact hash, open defect.

Unknowns

Unrepresented users/sites, production tails, dependency behavior, weak sample, missing UAT, unexecuted recovery, unresolved legal/security/privacy/safety decision.

Judgments

Risk tolerance, threshold, representativeness, acceptable residual condition, business value, authority disposition.

Preferences/pressure

Deadline, announcement, commercial expectation, stakeholder enthusiasm, sunk cost.

Pressure matters to the decision but cannot become technical evidence. [CLM-137]

Build a defensible recommendation

The FDE should state:

  • recommended disposition;
  • exact cohort/scope/duration;
  • evidence that supports it;
  • gaps/uncertainty and consequence;
  • alternatives and tradeoffs;
  • conditions/compensating controls;
  • success/guardrail/stop/recovery triggers;
  • decision authority and required signatures;
  • communications/support/command plan.

The FDE can recommend; designated customer authorities accept risk. [CLM-138]

Example:

Recommend delay for the safety-relevant cohort because qualified-review UAT and break-glass restoration are unexecuted. A reduced-scope, non-safety internal cohort is supportable for 48 hours if evidence binding, critical policy, latency/fallback, audit, support, and stop gates remain green. Release authority must sign; any cross-boundary denial anomaly, critical-quality failure, unreconciled intent beyond the owner limit, or unavailable recovery access stops exposure.

This is not indecision. It separates supported value from unsupported consequence.

Present real options in the meeting

Do not frame the decision as launch or fail. Present go for the exact supported cohort; conditional go under a legitimate exception; reduced scope that removes the unsupported state; delay until named evidence/authority arrives; and stop when a bounded safe path does not exist.

For each option, state value gained or delayed, users affected, residual uncertainty, operating cost, reversibility, evidence learned, and recovery. A reduced scope that excludes the hardest representative condition may produce little learning. A delay can preserve a safer migration window. A stop can be correct.

The designated authority chooses. Preserve the FDE recommendation, dissent, conditions, and trigger so future readers do not confuse the final decision with unanimous technical belief.

Choose rollout pattern by evidence need

Comparison of internal, shadow, canary, cohort, regional, and full rollout across user effect, evidence, comparability, representativeness, blast radius, reversibility, duration, support and recovery.
F15.2 - Rollout patterns buy different evidence. Pattern selection balances representativeness, blast radius, reversibility, support, and stop needs.

Internal

Real system under limited authorized internal users. Useful for operations/workflow rehearsal; may not represent customer users/data/dependencies.

Shadow

Observe/calculate without affecting user decision. Useful for comparison/capacity/data quality; carries data/privacy/cost risk and does not prove user behavior.

Canary

Partial/time-limited changed cohort versus control plus absolute criteria. Useful for release regression; weak if volume/cohort/quality signals cannot reveal target failure. [CLM-131]

Cohort or region

Selected users/sites/capability/region. Useful for representative learning and blast-radius control; requires legitimate selection, support, isolation, rollback/recovery.

Full

All intended exposure. Appropriate only when evidence/ownership/capacity/recovery justify it; still monitored and reversible where possible.

A small cohort is not automatically low consequence. One safety-critical site or irreversible external action can be high risk. [CLM-132]

Match pattern to the uncertainty

Choose shadow when the main unknown is data/quality/capacity and user effect must remain zero, while still governing data and cost. Choose canary when changed behavior can be compared meaningfully and stopped. Choose cohort/region when workflow, support, or environment representation matters. Use internal exposure for rehearsal, not as a substitute for customer evidence. Full exposure is not a learning shortcut.

Record why rejected patterns do not answer the decision. If Orchid’s uncertainty is whether technicians can evaluate evidence under intermittent connectivity, an office internal cohort is not representative. If safety-relevant recommendations cannot be safely shadowed because reviewers may treat outputs as advice, the shadow design needs stronger isolation or rejection.

Define the observation window

Record:

  • start/end and minimum volume/time;
  • expected workload/events/critical cases;
  • baseline/control and confounders;
  • signal freshness/coverage;
  • service, workflow, data, AI/policy, security, support, capacity, and cost criteria;
  • owner on watch and decision cadence;
  • criteria to widen, hold, reduce, stop, or recover.

Do not choose the window after seeing a convenient green period. A short quiet canary may prove little. Simulation strengthens evidence but cannot remove production-tail unknowns. [CLM-135]

Define minimum critical-case coverage as well as time and volume. A 48-hour window with no ambiguous equipment or degraded dependency cannot support those claims. If a needed case does not occur naturally and safe simulation is credible, inject it under controlled conditions and label the evidence type.

Precommit triggers and actions

Examples:

  • wrong equipment/evidence binding: stop affected path, preserve/reconcile, notify owner;
  • critical prohibited suggestion reaches review: disable bounded suggestion and investigate bundle/control;
  • tenant/region boundary anomaly: isolate and follow security/incident authority;
  • unknown intent beyond limit: stop new effects, reconcile;
  • low-connectivity cohort latency/failure: hold/reduce cohort, use fallback;
  • recovery access unavailable: stop widening, repair/rehearse;
  • support load exceeds capacity: pause ramp and fix workflow/training/support;
  • cost/capacity guardrail: restrict/defer/fallback.

Decide actions before impact. [CLM-140]

For every trigger bind detection, threshold/evidence state, affected scope, immediate action, owner/authority, communication, recovery path, and re-entry criterion. A dashboard threshold without an action is not a gate. A trigger decided after impact is a reaction, not a precommitment.

Make exceptions formal and expiring

An exception requires:

  • exact criterion/risk/consequence;
  • scope/cohort/duration;
  • alternatives/reduced scope considered;
  • compensating control and evidence;
  • monitoring/stop/recovery;
  • designated authority and delegation basis;
  • remediation owner/date;
  • expiry/review trigger;
  • communication to operators/support/users where needed.

Silence, attendance, or deadline is not acceptance. [CLM-136]

The companion makes an exception invalid without designated authority and expiry. It yields conditional go, never full go.

An exception cannot convert failed evidence into success. It says a named authority accepts a stated residual condition for a bounded scope under explicit controls. If authority, consequence, or expiry is missing, the row remains a gap.

Example:

Conditional go for internal non-safety users for 48 hours despite incomplete field-site connectivity evidence. Product release owner accepts reduced representativeness; low-connectivity sites are excluded, offline fallback remains available, support is staffed, and latency/fallback use is reviewed every four hours. Expiry is the end of the window. Any included-site loss of fallback or evidence-binding anomaly stops exposure. This exception does not permit the safety-relevant or low-connectivity cohort.

Prepare people and command

Readiness includes:

  • release/incident decision roles and contact schedule;
  • on-call/support staffing and access;
  • user/operator notice, eligibility, fallback, support path;
  • customer/security/privacy/safety/legal communications owners;
  • runbooks, dashboards/queries, reconciliation, feature controls;
  • change freeze/dependency/vendor coordination;
  • status cadence and evidence location.

Technical mechanics without operational ownership are not ready. [CLM-133] [CLM-139]

Resolve the Orchid deadline injection

At deadline, Orchid has passing synthetic evidence, a local recovery rehearsal, and an unexecuted customer recovery-access/control test.

It must not label the control met because documentation exists. The appropriate options are:

  1. delay all exposure until executed evidence;
  2. reduced-scope internal/non-safety cohort that does not rely on the missing control;
  3. formal conditional exception only if designated authority can legitimately accept the consequence and recovery conditions.

The companion test chooses delay for an unresolved gap even when the cohort is one site. Deadline pressure is recorded, not promoted.

Conduct the readiness review

Give participants the packet plus conflicting positions: a product stakeholder wants the announced date, operations wants a smaller cohort, security notes unexecuted recovery access, and technicians want the evidence display fixed.

The review must:

  1. confirm decision authority and the exact decision;
  2. separate facts, unknowns, judgments, and pressure;
  3. disposition every criterion;
  4. state the consequence of the material gap;
  5. compare go, conditional, reduced, delay, and stop;
  6. select cohort, pattern, window, and critical-case coverage;
  7. precommit widen, hold, stop, and recovery actions;
  8. staff command, support, and communications;
  9. record decision, conditions, dissent, and next trigger.

Pass when a reviewer can see why the exposure is supported and which evidence would change it. The exercise fails if the date determines evidence state or the FDE self-approves an exception.

Failure modes and repairs

Missing becomes not applicable

Repair: missing, stale, inconclusive, failed, and unexecuted evidence stays a gap until justified scope exclusion or formal decision. [CLM-130]

The small pilot is called safe

Repair: examine consequence, representativeness, external state, critical-case coverage, support, and recovery rather than user count. [CLM-132]

Canary ends when the chart is green

Repair: predefine time, volume, critical segments, absolute/comparative criteria, confounders, and decision actions. [CLM-131]

The meeting allows only go

Repair: present all five dispositions with consequence and opportunity cost. Readiness is a decision system, not launch theater. [CLM-134]

The FDE accepts adjacent risk

Repair: recommend and escalate; named customer authorities own formal exceptions and risk decisions. [CLM-138]

Triggers are invented during impact

Repair: bind detection, action, authority, communication, recovery, and re-entry before exposure. [CLM-140]

Operate the decision after exposure begins

A signed go record is not the end of readiness. During the observation window, preserve the selected cohort, bundle, evidence coverage, staffing, and conditions. Record every hold, exception, configuration change, support spike, missing signal, and departure from the plan.

At each cadence, choose one of:

  • widen: criteria and critical-case coverage support the next named cohort;
  • continue: the current window needs more evidence and conditions remain bounded;
  • hold: stop widening while a gap is investigated;
  • reduce: remove an unsupported cohort/capability without pretending full success;
  • recover: execute rollback, roll-forward, isolate, restore, or fallback by actual state;
  • stop: end exposure and preserve evidence when the path is not supportable.

The post-window record should compare expected and observed behavior, not merely say pilot complete. Include outcome/guardrail/adoption signals, critical segments, incidents/support, capacity/cost, recovery, limitations, confounders, and the next authority decision.

Keep the decision bundle stable enough to interpret

If code, configuration, model, prompt, index, policy, cohort, training, and support all change during one window, a favorable or unfavorable result becomes hard to interpret. Freeze material parts where possible and record necessary changes as new evidence versions.

An urgent correction may be necessary. Do not delay it for experimental purity. Contain harm, version the change, reset the claim/window if needed, and state what evidence is no longer comparable.

Distinguish rollout from migration cutover

A cohort feature can often be disabled. A data migration or external workflow cutover may mutate state that cannot be assigned back to a control group. For these changes, readiness emphasizes compatibility, capacity, reconciliation, recovery access, cutover command, and the last safe decision point.

Suppose a fictional analytics deployment replaces a nightly file with streaming events. A 5 percent consumer cohort does not limit upstream schema or backlog consequence if all events enter the new pipeline. The rollout unit must reflect the actual state boundary, not the UI audience. This non-AI transfer case changes the decision and prevents small cohort language from hiding shared infrastructure risk.

Write the final go/no-go record

Use this decision summary:

  • exact decision, scope, cohort, bundle, and window;
  • evidence packet version and critical claims;
  • facts, unknowns, judgments, and pressure;
  • met/gap/exception/reduced/block table;
  • options and consequences;
  • FDE recommendation and dissent;
  • designated authority and decision;
  • conditions, owners, expiry, and communications;
  • widen/continue/hold/reduce/recover/stop triggers;
  • post-window evidence and next decision.

Another reviewer should be able to understand why the authority chose the path even if they disagree. That auditability matters when later production evidence contradicts the original assumptions.

Complete OA-09

The readiness artifact includes:

  • criterion/evidence/disposition table;
  • facts/unknowns/assumptions/preferences;
  • options: go/conditional/reduced/delay/stop;
  • calibrated recommendation and consequence;
  • cohort/pattern/ramp/observation;
  • success/guardrail/stop/recovery triggers;
  • command/support/user/stakeholder communication;
  • designated authority/decision/signatures/conditions/expiry;
  • post-window decision and evidence update.

The companion adds five tests: missing evidence invalid, gap delays a small cohort, formal exception yields conditional go, reduced scope cannot become full go, and go/no-go records preserve unknowns/triggers/communications/recovery.

The Chapter 15 gate

Before exposure:

  • all required evidence is current or visibly gapped;
  • owner-accepted exceptions are formal/expiring;
  • readiness reviews operability/ownership proportional to risk;
  • cohort/pattern can reveal the target uncertainty;
  • observation has relative and absolute criteria;
  • command/support/user preparation and access are verified;
  • stop/recovery/communication triggers are precommitted;
  • named authority selects and signs the disposition;
  • FDE recommendation preserves alternatives and consequence;
  • companion readiness tests pass.

The gate passes when the decision record remains defensible even if the answer is delay or reduced scope.

Chapter 16, Stabilize Under Real Conditions, begins when production reality contradicts assumptions and the team must contain impact, communicate calibrated facts, restore service, correct the system, and prove stabilization.