Contract for an Outcome
Turn workflow evidence into measurable outcomes, guardrails, adoption signals, decision rights, assumptions, non-goals, and exit criteria.
A convincing prototype can be evidence of technical possibility and still be evidence against the deployment.
It may show that a model can summarize manuals, that an integration can return inventory, or that a user can complete a scripted task. It may also reveal that the team has not defined which workflow result matters, what can become worse, who may accept the consequence, how real use will be observed, or when the work should stop.
The outcome contract answers those questions before polish makes them harder to ask.
This is not a commercial contract. It does not set price, legal liability, or procurement terms. It is an engineering and decision artifact that binds a deployment hypothesis to measurable evidence, guardrails, authority, and exit conditions.
Official service guidance supports defining performance evidence early, establishing baselines, and monitoring benefit so a team can continue, change, or stop. Site reliability practice adds a compatible principle: user-relevant indicators and objectives are valuable when they lead to explicit action and tradeoffs. This chapter synthesizes those ideas into an FDE outcome contract. [CLM-021] [CLM-022] [CLM-024] [CLM-027]
Contract the change in work, not the feature
Chapter 3 reframed Orchid’s request. The bounded problem is not the absence of generated text. Evidence needed for selected service decisions is fragmented, inconsistently identified, variably available, and difficult to evaluate at the point of work.
A weak success statement is:
Launch an AI copilot to all technicians by the seasonal peak.
It names a mechanism, population, and date. It does not state value, guardrails, evidence, baseline, adoption, authority, or what to do if the mechanism fails.
A stronger outcome hypothesis is:
For a selected equipment family, region, and technician cohort, an evidence-assisted workflow may reduce the effort and delay required to reach a supported resolution or appropriate escalation, while preserving qualified safety decisions, evidence visibility, data/access boundaries, service reliability, cost control, and customer ownership.
The word may is essential. At this point Orchid has no measured baseline or accepted target in the fictional case. The outcome contract must record the measurements required rather than invent them.
The contract begins with six questions:
- Which operational result should change?
- For which users, cases, environment, and time boundary?
- Which evidence would show progress or failure?
- Which harm, degradation, or cost must not be traded away?
- Who has the authority to define, accept, stop, or revise each consequence?
- What action follows each evidence state?
If the final question has no answer, the metric is observation without governance.
Build a metric tree
A metric tree connects the intended workflow result to evidence at different distances from it.
Use five branches.
Outcome measures
Outcome measures describe the operational result users or the customer ultimately need. For Orchid candidates may include time to supported resolution or appropriate escalation, first-visit success under a precise definition, repeat service within a defined window, or proportion of selected cases reaching an evidence-complete decision.
These are candidates, not accepted metrics. The customer must define the operational meaning, population, baseline, target direction or threshold, and consequence.
Leading workflow measures
Leading measures describe steps plausibly connected to the outcome: time to identify equipment, time to gather required evidence, rate of unresolved identifier ambiguity, approval wait, inventory reconciliation time, or proportion of recommendations whose evidence is inspected before decision.
They are often available sooner than outcome measures. They can be useful for diagnosis. Improving them does not guarantee the outcome.
Adoption and use measures
Adoption asks whether the intended workflow is actually incorporated into representative work. Measures may include eligible cases in which the path is used, repeated use by the target cohort, completion versus abandonment, bypass reason, evidence-view behavior, and distribution across users/sites/case types.
Use is not value. Non-use is not automatically resistance. Use, satisfaction, dropout, and segmented workflow evidence are complementary signals. [CLM-023]
System and quality measures
These describe technical behavior required for the workflow: latency by site/cohort, error/timeout rate, identifier-match state, data freshness, recommendation/evaluation category, approval/audit completeness, retry/reconciliation state, and support burden.
System measures are usually causes, constraints, or diagnostic signals. Availability does not prove usefulness; model quality does not prove adoption.
Guardrail measures
Guardrails show whether apparent benefit is achieved by an unacceptable trade. Orchid guardrail candidates include safety-rule violation or unsafe recommendation categories, missed required escalation, incorrect equipment/part binding, privacy/access exception, reliability degradation, user effort moved into verification, support load, cost per eligible case, accessibility failure, and regional-data breach.
A guardrail can stop or reduce rollout even when the primary outcome improves.
Define every measure as a contract
A label such as “resolution time” is not a definition. A metric record should contain:
- stable metric ID and name;
- decision or promise served;
- exact numerator, denominator, unit, and event semantics;
- population, cohort, segment, exclusions, and window;
- start and end conditions;
- source systems and provenance;
- data quality/missingness limits;
- baseline method and version;
- target direction, threshold, or range, with owner;
- cadence and observation window;
- guardrail relationship;
- action when healthy, uncertain, or breached;
- accountable decision owner;
- known gaming/counterfactual risk;
- change history.
Consider “first-visit success.” It could mean no additional site visit within 30 days, issue resolved during the first arrival, or customer operation restored without escalation. It can be gamed by excluding difficult equipment, converting repeat visits into new tickets, or defining success before the equipment is tested. The useful definition must match the workflow state model and disclose exclusions.
For fictional Orchid, use placeholders such as BASELINE REQUIRED, TARGET OWNER REQUIRED, and WINDOW TO BE APPROVED. Do not replace missing evidence with a realistic-looking number.
Establish a baseline without freezing the past
A baseline describes the current measurement under a versioned definition and boundary. It is not “the number before launch” unless the collection semantics are comparable.
Baseline work should check:
- Does the current workflow emit the start/end events?
- Did definitions, policies, systems, or populations change during the window?
- Are difficult or abandoned cases missing?
- Can the measure be segmented by the target cohort without inappropriate data collection?
- Does measurement itself change behavior?
- Is the seasonal peak materially different?
- Which manual work is not captured?
- Can the same metric definition operate after the deployment?
If no credible baseline exists, say so. Options include a prospective baseline period, customer-reviewed sampling, a phased comparison, matched historical data with limitations, or a direction-only pilot claim. The decision owner must accept the weaker inference.
Benefit guidance emphasizes establishing a baseline and monitoring evidence that can support continuation, change, or stopping. It does not make the baseline causal proof. [CLM-022] [CLM-026]
Segment before the average reassures you
An aggregate can improve while a consequential segment worsens.
Orchid should consider segments such as equipment family/subfamily, region, site connectivity, technician experience, request channel, exception class, safety relevance, data completeness, and workflow route. Only collect/use segment attributes that are legitimate and necessary.
Suppose mean time to evidence improves. Low-connectivity sites may be slower. New technicians may abandon more often. High-risk equipment may have worse evidence retrieval. A model may perform well on common manuals and poorly on a minority equipment family.
The contract should identify precommitted critical segments whose results cannot be hidden in an average. It should also define minimum evidence volume or uncertainty treatment. A tiny segment should not produce a precise claim, but high consequence may still require a block or review.
Test every metric for gaming
Ask how a rational actor or system could improve the number without improving the intended result.
For each candidate metric:
- Can cases be excluded, transferred, cancelled, or relabeled?
- Can the start or end event move?
- Can difficult users disappear from the denominator?
- Can work move outside the measured system?
- Can speed improve by reducing evidence, quality, safety, or explanation?
- Can use be coerced or duplicated?
- Can an error become an escalation that another team absorbs?
- Which guardrail or audit would reveal the gaming path?
Orchid’s ticket closure metric is vulnerable. A difficult case can be reassigned, prematurely closed, split, or excluded from the selected category. The metric tree therefore treats closure as a process signal at most, not the production outcome.
Adoption is a diagnostic system
Adoption combines eligibility, access, repeated use, workflow fit, trust, and retained value.
Define the adoption funnel:
- eligible case occurred;
- user had access and the system was available;
- workflow path was offered or selected;
- user began the path;
- user reached a meaningful decision state;
- output/evidence was inspected or used as intended;
- user completed, escalated, or intentionally rejected the path;
- user returns for later eligible cases;
- support/bypass reason is recorded.
Each transition can fail for a different reason. A training campaign is not the default remedy.
A vendor customer story reports 98 percent active use by Morgan Stanley advisor teams. The figure is useful here only as an attributed example of adoption evidence; it is not independent proof of productivity, correctness, financial outcome, or customer benefit. [CLM-025]
This attribution discipline applies to Orchid too. Even if synthetic exercises later show a high use rate, the text must not convert it into a real customer outcome.
Connect service objectives to action
Technical objectives matter because a workflow depends on them. An SLO-like record should connect a user-relevant indicator to action and tradeoff, not choose a target because an industry example uses it. [CLM-024]
For example:
For eligible low-connectivity pilot requests, the evidence path must return a usable response or an explicit safe fallback within the approved latency objective for the stated proportion/window. A breach pauses cohort expansion and invokes the runbook; repeated breach triggers reduced scope or rollback under the named owner.
The objective remains incomplete until Orchid supplies the actual threshold and accepts the tradeoff. The book can teach the structure without inventing the number.
Calibrate the attribution claim
An outcome can move for many reasons: seasonality, staffing, policy, data cleanup, training, inventory availability, case mix, another system change, or simple measurement drift. The outcome contract should state what kind of inference the deployment can support.
Use an attribution ladder.
Descriptive
During the observation window, the contracted measures had these values for the stated cohort and definition.
This describes what was measured. It does not connect the change to the deployment.
Before/after association
The measure changed from the versioned baseline to the pilot window after the deployment was introduced.
This establishes timing and direction under the definitions. It remains vulnerable to case mix, seasonality, concurrent changes, and regression to the mean.
Comparative association
The exposed cohort changed differently from a defined non-exposed or earlier cohort under stated comparability assumptions.
This strengthens the evidence if assignment, populations, and measurement are credible. It does not automatically remove selection or spillover.
Plausible contribution
The deployment changed specific leading workflow states in the expected direction; outcome/guardrail movement was consistent with the mechanism; alternative explanations were investigated and remain bounded as stated.
This combines process evidence with outcome evidence. It can support a professional decision without claiming experimental causality.
Causal estimate
The study design supports a defined causal effect estimate under its assumptions.
This requires appropriate experimental or quasi-experimental design, sufficient evidence, and specialist review where consequence demands it. Most FDE pilots should not casually claim this level.
The contract chooses the intended claim level before the results are known. That reduces the temptation to promote a favorable association into causality after launch.
For Orchid, a bounded cohort before/after comparison may be feasible, but the seasonal peak, equipment mix, technician experience, and concurrent process changes could matter. The initial contract should plan to report descriptive and associated change plus mechanism/limitations unless stronger design is approved.
Design the measure from events and states
The workflow state model from Chapter 3 is the best defense against metric ambiguity.
Suppose Orchid selects “time to supported decision.” Define it from state transitions:
- Start: request first enters
triagedwith sufficient identity to enter the selected workflow population. - End: request enters
supported-resolution,qualified-escalation, or another approved terminal decision state. - Paused time: decide whether customer/site unavailable time is included, excluded, or reported separately.
- Reopened work: define whether reopening continues the same episode and for how long.
- Transfers: preserve the original episode if the same operational result remains unresolved.
- Cancelled cases: classify by reason and never exclude silently.
- Missing end: retain as unresolved/censored under the defined analysis rather than dropping it.
- Version: state contract and metric definition versions remain linked.
Now the team can test event emission and query behavior before production. A dashboard label cannot replace this definition.
For “evidence-complete decision,” the contract must name required evidence classes and who defines completeness. A model citation count is not sufficient. The appropriate qualified owner may require equipment identity, current manual version, observed condition, inventory state, policy check, and approval evidence for a particular decision class.
The FDE can implement and verify the measurement. The domain owner defines the criterion where domain authority is required.
Pair rate and burden
A percentage can improve by adding hidden work. If evidence-complete decisions rise because technicians spend twice as long verifying output, the outcome may not improve.
Pair selected rates with burden evidence:
- user time and steps;
- context switches or duplicate entry;
- support contacts;
- exception queue age;
- reviewer/approver load;
- correction/reconciliation work;
- latency and abandonment;
- qualitative evidence about trust and fit.
Not every burden requires continuous tracking. The contract identifies which burden could reverse the benefit and how it will be sampled or monitored.
Measure the denominator before celebrating the numerator
If 80 of 100 surfaced recommendations are accepted, the acceptance rate is 80 percent. If only 100 of 1,000 eligible cases surfaced because difficult cases were filtered or failed earlier, the workflow story is different.
Maintain population flow:
eligible -> accessible/available -> offered -> started -> evidence produced -> decision -> completed/escalated/rejected -> retained outcome
Report loss and reason at each transition. Unknown reason is a meaningful category until resolved.
Turn evidence states into decisions
An outcome contract should include a decision table before results exist.
| Evidence state | Example interpretation | Permitted action |
|---|---|---|
| outcome improves, guardrails healthy, adoption representative, operations ready | hypothesis supported within limits | continue/ramp under owner approval |
| leading measure improves, outcome window incomplete | mechanism signal only | hold scope; continue observation |
| outcome improves, critical segment/guardrail fails | benefit not acceptable for current path | stop/reduce/repair; no aggregate override |
| use low, technical health good | cause unknown | diagnose workflow/access/trust; do not force rollout |
| outcome unchanged, burden/support rises | hypothesis not supported | stop or redesign |
| evidence quality insufficient | claim cannot be made | improve measurement, reduce claim, delay, or stop |
| benefit observed, ownership/recovery unproven | customer outcome not yet supportable | do not exit/ramp; complete operations/handoff evidence |
The table prevents result interpretation from becoming a negotiation after the deadline. It also makes a negative or inconclusive result useful. A stopped deployment can be a successful decision if it prevents unjustified exposure or investment.
Resolve competing outcomes explicitly
Stakeholders can agree that a deployment should “improve service” while optimizing different consequences.
For Orchid:
- service leadership values throughput and seasonal capacity;
- technicians value safe, evidence-supported resolution;
- customers value equipment availability and trustworthy communication;
- inventory values stock efficiency and correct reservation;
- security/privacy owners value bounded access and handling;
- support values diagnosable failure and manageable demand;
- product teams value reusable capability;
- finance may value cost predictability.
The outcome contract does not hide these interests inside one weighted score. Weighted scores can make serious guardrail failure appear acceptable because another benefit is large.
Use a hierarchy:
- prohibited or formally constrained outcomes;
- minimum guardrail conditions;
- primary bounded outcome;
- secondary benefits and burdens;
- diagnostic/leading measures.
Where two valid outcomes conflict above a threshold, route the tradeoff to the named authority. The FDE documents alternatives, evidence, reversibility, and consequence. A metric formula should not silently accept risk.
Version the outcome contract
The contract will change as the scope and system become concrete. Version changes should name:
- evidence that triggered the change;
- affected metric/definition/population/owner;
- whether baseline comparability breaks;
- which tests, dashboards, rollout triggers, or claims change;
- who approved the change;
- effective date and prior-version disposition.
Never adjust a threshold after seeing an unfavorable result without recording the change. A legitimate definition correction is possible, but the old result and reason remain visible. Otherwise the team teaches the metric to pass.
Build a guardrail ladder
Not every guardrail uses the same response. Classify the control action:
- prevent: prohibited action cannot execute;
- require approval: qualified identity must accept before action;
- detect and block: evidence state triggers a stop/fallback;
- detect and escalate: action pauses for an owner;
- monitor and review: signal enters a scheduled decision;
- limit exposure: cohort, capability, rate, or region remains bounded;
- stop/rollback: threshold or event ends exposure;
- record limitation: residual uncertainty is visible and accepted by the authorized owner.
Safety-relevant eligibility may be prevented or require qualified approval. Low-confidence evidence matching may block recommendation. Latency degradation may limit exposure. Support demand may trigger review. The ladder connects the metric to an actual production decision.
Separate assumptions, constraints, and decisions
An outcome contract includes an assumption register.
- Assumption: unverified proposition with owner, consequence, and validation path. Example: selected-family manuals are sufficiently current for evidence retrieval.
- Constraint: real boundary shaping the solution. Example: selected regional data cannot leave an approved boundary.
- Decision: accepted choice with evidence, alternatives, authority, and revisit trigger. Example: pilot excludes autonomous parts ordering.
Do not promote an assumption into a constraint because it is inconvenient to test. Do not call a preference a requirement. Do not treat a constraint as immutable if the accountable owner can change it.
Map decision rights
For each outcome, guardrail, scope, and gate decision, name four actions:
- recommend: assembles evidence/options and proposes a path;
- approve/accept: holds authority for the consequence;
- execute: performs the change or operation;
- inform: needs timely decision/result information.
The FDE can recommend the technical path and maintain the evidence chain. Orchid’s service owner defines/accepts business outcome consequences. Qualified safety owners define approval rules. Security/privacy/regional owners handle their formal decisions. The customer’s rollout authority accepts go/no-go consequences. Execution can belong to different engineering or operations teams.
An owner’s absence is a blocker, not permission for the FDE to inherit authority.
Write non-goals and rejection criteria
Non-goals prevent the outcome statement from becoming an unlimited promise. For the initial Orchid contract:
- not all equipment families, regions, or channels;
- not automatic safety-relevant action;
- not enterprise identifier remediation;
- not replacement of customer safety/security/privacy/support ownership;
- not proof that AI is superior to deterministic or non-AI alternatives;
- not guaranteed causal attribution from a bounded pilot;
- not platformization of customer-specific code.
Rejection criteria state when a technically impressive result is insufficient:
- output lacks evidence needed for qualified judgment;
- aggregate quality hides a critical segment failure;
- use requires unsafe access or violates a data boundary;
- the workflow adds more verification/support burden than it removes;
- no credible recovery or support owner exists;
- the outcome cannot be measured under an accepted limitation;
- the first safe path cannot be bounded before the seasonal peak.
Define entrance, exit, and stop criteria
An outcome contract governs several gates.
Discovery exit
The current workflow, material exceptions, outcome hypothesis, stakeholders, access, and evidence gaps are sufficient to define measures and safe scope.
Pilot entrance
The selected cohort/path has verified behavior, controls, operability, support/user preparation, rollback/stop triggers, and authorized risk disposition.
Pilot exit
The agreed outcome, guardrail, adoption, and operating evidence has been observed over the accepted window, with limitations and exceptions dispositioned.
Stabilization exit
Material incidents/corrective actions are addressed, signals and support ownership operate, representative users can perform the path, and residual risks are accepted.
Engagement exit
The customer demonstrates operation/change/recovery, access is transferred or revoked, outcome review is complete, and reusable work has an isolation-aware disposition.
Stop criteria
Stop or reduce the path when a prohibited event occurs, a critical guardrail breaches, required approval disappears, evidence cannot support the claim, support/recovery is not viable, or new evidence invalidates the outcome hypothesis.
An outcome contract should make continue, change, reduce, delay, or stop legitimate outcomes. [CLM-026]
The OA-03 Orchid outcome contract
The fictional artifact now contains the structure below. Values remain intentionally unresolved until a hypothetical Orchid owner supplies or approves them.
Population. Selected equipment family, Region West, named pilot technician cohort, eligible service-request definition pending.
Outcome candidates. Time to supported resolution/appropriate escalation; evidence-complete decision rate; first-visit success under an approved definition.
Leading candidates. Equipment-match time/state; evidence-gathering time; approval wait; reconciliation wait.
Adoption candidates. Eligible cases offered/started/completed/escalated/rejected; repeated use; bypass reason; evidence-view behavior; segment coverage.
System/quality candidates. Latency by site; dependency/timeout/retry state; identifier/data quality; recommendation error class; audit/approval completeness; support demand/cost proxy.
Guardrail candidates. Safety-rule/approval violation; incorrect match/part; missed escalation; privacy/access exception; reliability; accessibility; low-connectivity degradation; verification burden; support load; cost.
Required evidence fields. BASELINE REQUIRED, DEFINITION OWNER REQUIRED, TARGET/THRESHOLD REQUIRED, OBSERVATION WINDOW REQUIRED, ACTION/OWNER REQUIRED.
Decision rights. FDE recommends/maintains evidence. Service, safety, security, privacy/regional, operations/support, and rollout owners approve within named boundaries.
Non-goals/rejection/exit/stop. As defined above and versioned as discovery/scope evidence changes.
The artifact is useful before it has numbers because it identifies exactly which numbers, definitions, owners, and actions are missing. It is not ready for go/no-go until those gaps are resolved or formally dispositioned.
Outcome-contract failure modes
Vanity success
Logins, prompts, outputs, or demos become the headline.
Repair: trace activity through the adoption funnel to workflow outcome and guardrails.
Aggregate comfort
The mean improves while a critical region, user, or equipment segment fails.
Repair: precommit material segments and consequence-based review.
Target without baseline
A realistic-looking target is chosen before event semantics and current performance are known.
Repair: define the measurement and baseline method; mark target authority/evidence as missing.
Guardrail without action
A risk metric is monitored but no threshold, owner, or stop path exists.
Repair: assign a control response and decision right.
Causal overclaim
Observed improvement during a pilot is attributed entirely to the deployment.
Repair: state design, comparison, confounders, counterfactual limits, and calibrated claim.
FDE acceptance
The FDE signs off a safety, privacy, security, business, or support consequence because the owner is absent.
Repair: escalate, reduce, delay, or stop; keep the missing decision visible.
Conduct the outcome-contract exercise
Build the fictional Orchid metric tree, then fully specify one outcome, two leading workflow measures, one adoption measure, one system-quality measure, and three guardrails. For each define population, event/state, source, formula, segment, window, baseline, threshold authority, gaming path, missing-data behavior, owner, and decision trigger.
Add a decision-rights map, assumptions with validation, non-goals, rejection criteria, and entrance/exit/stop conditions. Inject a favorable average that hides one low-connectivity cohort and a faster closure rate caused by transferred cases. Repair the contract without changing definitions after seeing results.
Pass when the contract can reject a convincing but useless or harmful deployment, preserves attribution limits, and tells the designated owner what action each evidence state supports.
The Chapter 4 gate
Before Chapter 5 scopes the path, OA-03 must contain:
- bounded population/workflow outcome hypothesis;
- metric tree with definitions and data feasibility;
- baseline plan and explicit missing evidence;
- critical segments;
- adoption funnel and bypass evidence;
- guardrails linked to controls/actions;
- assumptions, constraints, non-goals, and gaming risks;
- recommend/approve/execute/inform decision map;
- discovery, pilot, stabilization, engagement, and stop criteria;
- attribution limitations and change history.
The gate does not require invented certainty. It requires that uncertainty has a name, consequence, validation path, owner, and effect on the decision.
Chapter 5, Scope the First Safe Production Path, will use this contract to choose the least exposure that can produce decision-quality evidence across the real user, data, integration, control, environment, and operating boundaries.