Write the Behavior Contract
Specify required, uncertain, abstaining, escalating, degraded, and prohibited behavior with explicit evidence and authority.
“Helpful and safe” is not a specification
Imagine that Patchwork’s team writes one requirement for its part finder: “Give helpful, accurate, and safe recommendations.” Everyone agrees. Nobody can implement or test it.
Does helpful mean broad recall or a short list? Does accurate mean textually relevant, dimensionally compatible, or guaranteed to fit? Does safe mean excluding high-consequence categories, protecting private data, avoiding prohibited claims, or requiring specialist approval? What should happen when a buyer omits shaft diameter? What if two catalog fields conflict? May the system draft a seller question? May it send one? Who decides whether a domain exception is acceptable?
The vague requirement creates false alignment. Product imagines a useful candidate set. Engineering imagines a ranking metric. Design imagines a confident explanation. Trust imagines conservative claims. A domain reviewer imagines verified measurements. The first integrated behavior is where these private meanings collide.
A behavior contract turns product intent into observable states, evidence requirements, and decision rights. It does not eliminate uncertainty. It specifies how the product behaves because uncertainty exists.
Trustworthiness is contextual. The same output can be acceptable for low-consequence exploration and unacceptable for an irreversible or high-consequence decision. The contract must connect intended use, affected people, consequence, and the organization’s risk tolerance. [CLM-011]
What a behavior contract is
A behavior contract is a versioned agreement about the combined application’s externally meaningful behavior. It covers learned components, deterministic logic, data/context, interfaces, controls, human roles, and operating states.
It is not:
- a provider’s model behavior policy;
- a prompt;
- an output JSON schema;
- an exhaustive list of correct answers;
- legal or marketplace policy drafted by an engineer;
- a guarantee that a probabilistic component will never fail;
- a release approval.
Provider behavior specifications can reveal how a model is intended to resolve instructions. Structured-output features can constrain syntax. Neither defines Patchwork’s evidence semantics, user promise, authority, or safe effect. A perfectly valid compatible: true field can still be false, unauthorized, or unsupported. [CLM-014]
The contract instead creates a surface against which mechanisms, interfaces, evaluations, controls, and release evidence can be compared.
Contract anatomy
A complete clause has more than a desired result. Use this structure:
| Field | Question |
|---|---|
| ID and version | Which exact clause is this, and when did it change? |
| Trigger and segment | In which task state and for which users/items does it apply? |
| Required behavior | What must the combined product do? |
| Allowed variation | Which different outputs remain acceptable? |
| Unacceptable behavior | What failure violates the clause? |
| Evidence requirement | What input or context must exist before the behavior is allowed? |
| Uncertainty behavior | When must it clarify, qualify, abstain, or escalate? |
| User interaction | What does the user see, correct, or choose? |
| Authority | Who supplies judgment, who decides, and who may execute? |
| Evaluation | Which cases and criteria test the clause? |
| Release implication | Does failure block, narrow, degrade, or merely inform release? |
| Change trigger | Which model, data, policy, interface, or operating change requires replay? |
For example:
PF-B07 v0.1: For an in-scope seal query with an identified equipment family but a missing required shaft diameter, Patchwork must not present any candidate as compatible. It must ask for the measurement using domain-approved instructions or offer ordinary search without a compatibility claim. Tests cover missing, malformed, contradictory, and unitless values. Failure blocks the assisted-compatibility state for that segment.
The clause does not specify an exact sentence. It specifies criteria, evidence, prohibited interpretation, fallback, tests, and release consequence.
The behavior state model
A product with uncertain components needs more than success and error. Six states make Patchwork’s behavior inspectable.
Respond
Required evidence exists, the request is in scope, controls pass, and the product can return an allowed result. The response still names evidence and limitations. Responding does not imply certainty.
Clarify
The request may become answerable if the user supplies a bounded missing fact. The system asks only for information tied to a clause. It does not create an endless conversational burden or collect data speculatively.
Abstain
The system cannot justify the requested behavior from available evidence. It states what is missing and provides a safe next path. Abstention is a product behavior, not an exception stack trace.
Escalate
A competent named person or authority must decide. Escalation includes what evidence is sent, expected response time, what happens while waiting, and what the reviewer may decide.
Degraded mode
A dependency, budget, or control prevents normal behavior, but a bounded alternative remains safe. Patchwork might disable assisted compatibility while keeping ordinary catalog search and filters available.
Prohibited
The product must not perform or facilitate the action in the current contract. Patchwork cannot claim guaranteed fit, contact sellers automatically, or execute purchases. A user request does not override the boundary.

The state names are durable; thresholds are contextual. A low-consequence discovery product may respond with broader candidates. A high-consequence task may require specialist escalation or remain prohibited. Do not copy thresholds across consequences merely because the interface looks similar.
Specify criteria, not a fictional exact answer
Some software behavior can be expressed as exact input-output examples. Many AI-assisted tasks admit several acceptable outputs. A behavior contract can still be rigorous by specifying criteria.
A Patchwork explanation may vary in wording while satisfying all of these criteria:
- every factual attribute shown comes from a named listing or approved catalog source;
- matched and unresolved attributes are distinct;
- no phrase converts similarity into guaranteed compatibility;
- a user can inspect the evidence and correct their measurements;
- missing critical constraints trigger clarification or abstention;
- prose does not conceal a deterministic exclusion.
The evaluation can use examples, counterexamples, deterministic checks, and competent judgment. It need not pretend there is one canonical paragraph.
Criteria should include ordinary variation. If every test expects the same wording, the team may overfit a prompt without verifying semantics. They should also include adversarial and accidental variation: swapped units, ambiguous abbreviations, contradictory listing text, prompt-like seller content, empty retrieval, timeouts, duplicate candidates, and user corrections.
Uncertainty is not one number
Teams often ask for a confidence threshold as if uncertainty were a single calibrated fact. Different uncertainty sources demand different behavior:
- input uncertainty: the user has not supplied a required measurement;
- semantic uncertainty: a catalog field has ambiguous meaning;
- retrieval uncertainty: evidence may be absent because the index is stale or incomplete;
- model uncertainty: the learned mechanism’s score is not decisive;
- judgment uncertainty: competent reviewers disagree;
- system uncertainty: a dependency version or response cannot be verified;
- authority uncertainty: nobody present holds the required decision right.
A numeric score can support one part of this system. It cannot summarize all seven.
Research on calibration distinguishes accuracy from confidence reliability, and calibration quality depends on model and data conditions. Selective classification formalizes a coverage-risk tradeoff: a system can decline some cases to reduce error among the cases it accepts. Those results motivate measurement, not a universal threshold. Application abstention, user effort, and escalation still require product design. [CLM-013]
For Patchwork, a risk-coverage report might ask: as the candidate-ranking acceptance threshold rises, what fraction of in-scope cases receive assisted results, and what consequential error remains among them? Segment the curve by catalog family, attribute completeness, unit quality, and other material conditions. A favorable aggregate can hide a segment with dangerous acceptance behavior.
Raw score labels should not be translated into “90 percent compatible” unless the probability semantics are valid for that outcome and segment. Often a safer interface uses evidence states: required attributes matched, one critical attribute unresolved, or conflicting catalog evidence.
Clarification, abstention, and escalation have costs
Conservative behavior is not automatically good behavior. A system that abstains on every difficult case is safe only in a narrow technical sense; it may be useless, create unequal burden, or move risky decisions into unofficial channels.
Measure:
- coverage by task segment;
- error among accepted cases;
- clarification completion and abandonment;
- time and cognitive cost;
- escalation volume and response time;
- reviewer disagreement;
- fallback success;
- affected-party burden.
Then connect the measurements to consequence. If Patchwork asks for a dimension that most buyers cannot obtain, the task boundary may be wrong. Lowering the threshold is not the only response. The team could narrow to equipment families with accessible measurements, improve instructions, rely on part numbers, or stop compatibility assistance.
Every abstention should provide a next action appropriate to the state. “I cannot help” is insufficient when ordinary search can continue. Every clarification should explain why the fact matters. Every escalation should preserve what was already learned so the reviewer does not restart the task.
Users must be able to form and repair expectations
Human-AI interaction guidance emphasizes communicating capability, supporting correction, and handling failure over time. Scenario-based integration testing matters because a component can look acceptable in isolation while the human-system interaction fails. [CLM-012]
Patchwork’s contract therefore includes interface behavior:
- label the feature as candidate discovery, not a fit guarantee;
- show which user and listing attributes were used;
- separate matches, mismatches, and unresolved facts;
- let the user correct measurements without starting over;
- preserve ordinary search and filters;
- explain why a clarification is needed;
- reveal when assisted behavior is unavailable or degraded;
- never hide a deterministic exclusion behind generated prose.
Expectation setting is not a disclaimer placed below a confident recommendation. The primary label, visual hierarchy, state transition, and available action must agree. A red warning after a green compatibility badge does not produce calibrated trust.
Correction is also evidence. If users repeatedly change the same inferred field, the system may have a semantic defect. Do not treat every correction as training data automatically. It has purpose, provenance, permission, and possible strategic behavior that require separate review.
Human review is not a control until it is designed
“A human is in the loop” says only that a person appears somewhere. It does not say what they know, see, decide, or can prevent.
Design review with five questions:
- Competence: What expertise is required for this judgment?
- Evidence: What context, provenance, alternatives, and uncertainty does the reviewer receive?
- Time: Can the review happen before consequence, at the expected volume and response time?
- Reversibility: What can the reviewer prevent or correct?
- Decision right: Is the reviewer advising, approving, authorizing, or executing?
Domain-specialist involvement in high-consequence dataset work illustrates why expertise and provenance must be explicit. It does not authorize a generic transfer from clinical cases to marketplace decisions. The durable principle is to name the qualified role and its authority. [CLM-015]

In Patchwork, a buyer may correct their own measurements. A catalog specialist may determine whether two attributes have compatible semantics. A trust-policy owner may define prohibited marketplace claims. A product and risk authority may approve bounded release. The Applied AI Engineer implements the states and produces evidence; the role does not inherit those decisions. [CLM-015]
If a reviewer is unavailable, the system needs a state. It may remain abstaining, use degraded mode, or block the action. “Manual review” without a queue, response objective, competence, evidence packet, and fallback is deferred failure.
Degraded behavior belongs in the contract
Normal-path quality attracts attention; dependency failure determines whether the product remains honest.
Patchwork may depend on a catalog index, attribute service, unit normalizer, retrieval/ranking service, explanation component, policy rules, and telemetry. Define behavior for each failure.
Examples:
- If explanation generation fails but candidate evidence is intact, show structured evidence without generated prose.
- If ranking fails but filtered retrieval works, show an unordered candidate set labeled accordingly.
- If the authoritative attribute service is stale beyond its allowed version, disable compatibility assistance.
- If policy rules cannot load, fail closed for assisted claims while retaining ordinary search.
- If telemetry fails, the release gate may disable the experimental path because required operating evidence cannot be produced.
Degradation should not silently change the user promise. A badge must not remain green when the evidence service is unavailable. The interface should say which behavior is unavailable and what remains usable.
Prohibited behavior must be concrete
“Avoid harm” is not enough. A prohibition names an action, representation, or effect.
Patchwork’s current prohibited set includes:
- state or imply guaranteed compatibility;
- override a known deterministic incompatibility;
- fabricate a listing attribute, source, measurement, or seller statement;
- present an inferred measurement as seller-provided fact;
- use a private message or buyer history without approved purpose and access;
- draft or send content that pressures a seller to make an unsupported claim;
- contact a seller automatically;
- initiate or execute purchase;
- operate on excluded high-consequence categories;
- conceal degraded evidence behind normal-state presentation.
Each prohibition needs a detection and response design. Some can be prevented deterministically. Some require evaluation. Some require both. A prompt saying “never fabricate” is an intention, not a complete control.
Patchwork behavior and authority contract v0.1
PF-02 contains fifteen initial clauses. They are constructed product inputs, not claims about a real marketplace.
| Clause | Required behavior | Evidence and release implication |
|---|---|---|
| PF-B01 In-scope entry | Accept only named low-consequence catalog families in v0.1 | Unrecognized/excluded family routes to ordinary search; violation blocks release |
| PF-B02 Evidence display | Show source listing fields used for every candidate | Missing provenance removes candidate from assisted set |
| PF-B03 Ordinary variation | Explanation wording may vary while criteria remain satisfied | Semantic criteria and counterexamples, not exact string match |
| PF-B04 Known exclusion | A deterministic incompatibility always overrides rank/generation | Control test is release-blocking |
| PF-B05 Missing measurement | Ask for a domain-required fact and explain why | No compatibility language before value exists |
| PF-B06 Unit ambiguity | Do not infer a missing unit when alternatives alter fit | Clarify or abstain; critical segment |
| PF-B07 Conflicting evidence | Label conflict and withhold compatibility implication | Domain escalation only if queue and authority exist |
| PF-B08 Empty retrieval | State that no supported candidates were found | Never generate a candidate absent from authorized catalog evidence |
| PF-B09 Uncertainty | Express matched, unresolved, and conflicting constraints | Raw model score is not a user-facing fit probability |
| PF-B10 User correction | Permit correction and recompute from recorded values | Correction remains visible in trace; no automatic reuse as training data |
| PF-B11 Degraded explanation | Fall back to structured evidence if generation fails | Candidate logic must remain unchanged |
| PF-B12 Degraded authority data | Disable assisted compatibility if authoritative semantics are stale | Ordinary search remains available and visibly distinct |
| PF-B13 Seller question draft | May draft a bounded question only after explicit user request | Draft is editable; it cannot claim facts or be sent automatically |
| PF-B14 Prohibited effects | No seller contact, purchase, guaranteed fit, or excluded category | Deterministic capability boundary; violation is critical |
| PF-B15 Change replay | Catalog schema, rule, model, provider, prompt, index, or interface-semantic change triggers mapped tests | No silent promotion of a changed evidence version |
Authority map
| Decision | Prepares evidence | Reviews | Authorizes |
|---|---|---|---|
| Required measurement semantics | Applied AI + catalog engineering | Catalog specialist | Catalog domain owner |
| Prohibited marketplace claims | Applied AI + trust engineering | Trust/policy specialist | Designated policy/legal authority |
| New personal-data purpose | Product + Applied AI | Privacy/security specialists | Designated privacy/product authority |
| Behavior evaluation disposition | Applied AI Engineer | Domain, product, and quality reviewers | Release authority for bounded cohort |
| Residual release risk | Applied AI recommends | Named specialist owners | Designated product/risk authority |
| Seller contact or purchase | Not in implementation scope | Not applicable | Requires a future contract and independent authorization |
The tables do not settle policy. They state which missing decision blocks implementation.
The machine-readable contract
A contract should be readable by people and inspectable by tools. The companion begins with a provider-neutral JSON schema and a Patchwork v0.1 instance. The schema requires:
- contract identity and version;
- explicitly synthetic case status;
- scope and prohibited effects;
- state definitions;
- clauses with triggers, required behavior, evidence, failure disposition, authority, and change triggers;
- a named ordinary-search fallback.
Machine readability supports linting, traceability, and later evaluation generation. It does not create semantic correctness. The local companion test can prove that all required fields exist, IDs are unique, every clause names authority, and prohibited actions are present. It cannot prove that the domain rule is valid or that release is acceptable. [CLM-014]
This separation is essential. Use deterministic validation for structural claims and competent review or task evaluation for semantic claims.
Red-team the contract before the model
Contract red-teaming asks how the product could satisfy the words while violating the intent.
Fluent unsupported candidate
The candidate exists, the output schema is valid, and the explanation cites listing text. A required dimension is absent.
Contract response: PF-B05 or PF-B06 prevents compatibility language and requires clarification or abstention.
Green badge, cautious footer
The main card says “compatible” while a footer says “verify before purchase.”
Contract response: PF-B09 and PF-B14 govern primary representation; a disclaimer cannot repair a prohibited implication.
Safe model, unsafe tool
The model drafts a reasonable seller question, and orchestration sends it automatically.
Contract response: PF-B13 permits an editable draft only; PF-B14 prohibits contact. Tool authorization is a deterministic application boundary.
Human rubber stamp
A support agent receives hundreds of escalations with no source evidence and a ten-second response expectation.
Contract response: escalation is unavailable because competence, evidence, and time are not credible. The system abstains or narrows coverage.
Aggregate threshold hides a segment
Overall accepted-case error is low, but unitless legacy listings have high error.
Contract response: PF-B06 is a critical segment with its own disposition. Aggregate evidence cannot waive it.
Provider upgrade preserves syntax
A new model version returns the same JSON fields but changes which candidates it describes persuasively.
Contract response: PF-B15 replays semantic and interaction evaluations. Schema conformance alone is insufficient.
Degraded service looks normal
The attribute service times out, and cached candidates still show a normal compatibility badge.
Contract response: PF-B12 changes product state and presentation. Stale authority evidence cannot masquerade as current.
Contract red-teaming is cheaper than discovering these ambiguities after an architecture has made them difficult to correct.
Release implications belong beside the clause
Not every failure has the same disposition.
- Block: A prohibited effect occurs or a critical clause lacks evidence.
- Narrow: Exclude a segment, input state, or behavior while preserving a justified subset.
- Degrade: Use a pre-specified lower-capability path with an honest user promise.
- Warn and monitor: A non-critical limitation remains inside authorized tolerance with signals and an owner.
- Inform: Evidence improves understanding but does not change exposure.
The Applied AI Engineer recommends a technical disposition. The named release and risk authorities decide whether the evidence and residual consequence are acceptable. Formal authority does not emerge from a green test suite. [CLM-011]
Turn clauses into an evaluation design
A contract becomes useful when its clauses can generate cases and judgments. Do not wait until the mechanism is built to decide what counts as acceptable behavior.
Create a clause coverage matrix
For every clause, list ordinary, boundary, counterexample, dependency-failure, correction, and change cases. The matrix should show which evaluation method addresses each claim.
| Clause kind | Example case | Judgment method | Evidence limitation |
|---|---|---|---|
| Deterministic exclusion | Shaft diameter outside approved tolerance | Exact control assertion | Domain rule must itself be authorized |
| Evidence provenance | Candidate attribute shown with source listing/version | Schema and trace assertion | Presence does not prove source truth |
| Explanation criteria | Wording varies without stronger compatibility implication | Guided semantic review + deterministic phrase checks | Review disagreement and unseen variation remain |
| Clarification | Required measurement missing | State-transition test + interaction review | Completion in a test may not predict real burden |
| Abstention | Conflicting critical attributes | State and message criteria | Threshold must be evaluated by segment |
| Degraded mode | Explanation dependency times out | Failure injection | Test cannot cover every compound outage |
| Authority | Catalog conflict enters review | Queue, evidence-packet, and permission test | A functioning queue does not validate the final domain judgment |
| Prohibited effect | Tool request attempts seller contact | Capability/permission test | Other indirect effect paths still require threat analysis |
The matrix prevents one evaluator from becoming universal. Deterministic tests are excellent for exact exclusions, permissions, and structure. They cannot judge every explanation. Domain reviewers can judge semantics but may disagree or operate at limited capacity. Model-based graders may help scale a bounded criterion, but they introduce their own behavior and must be calibrated later. Chapter 11 will build that judgment system; this chapter creates the contract it must serve.
Derive cases from states
For each ordinary respond case, create neighboring cases that should clarify, abstain, escalate, degrade, or prohibit. If a case with 16 mm is accepted, test 16 with a missing unit, conflicting 16 mm and 5/8 inch sources, a stale source, an excluded category, and a user correction from 16 to 18 mm.
This technique reveals whether state transitions depend on real evidence or on superficial phrasing. It also exposes discontinuities. A one-character unit change may need a different state even when text similarity remains high.
Test the primary representation
Review what the user sees first, not only the complete response. A green badge can dominate a cautious explanation. A ranked first position can imply endorsement even without a label. A disabled correction control can make uncertainty effectively invisible.
Scenario-based human-AI testing should include sequence: expectation before use, result interpretation, correction, repeated use, failure, and recovery. The combined interaction can fail even when every static screen appears defensible. [CLM-012]
Include negative capability
Evaluate whether the system can refrain from behavior. Test that it does not invent a candidate after empty retrieval, does not infer a missing unit, does not reveal restricted data, does not execute a seller-contact tool, and does not preserve a normal badge in degraded mode.
Negative capability needs observability. A test should distinguish “the model happened not to do it” from “the application boundary made it impossible.” Prefer deterministic enforcement for prohibited effects.
Attach a disposition
Every failed case must map to block, narrow, degrade, warn-monitor, or inform. Without disposition, evaluation produces a dashboard rather than a decision. A critical PF-B14 violation blocks. An explanation-style preference may inform. A segment-specific abstention burden may narrow. The mapping must be reviewed before a release result creates pressure.
Maintain the contract through change
Behavior contracts decay when versions are implicit. Treat the contract as an artifact with lifecycle discipline.
Version meaningful semantics
Increment the contract when the user promise, included segment, state threshold, prohibited effect, authority, evidence requirement, or release disposition changes. Editorial clarification can be recorded without pretending behavior changed. Preserve a change log explaining why the new version exists.
PF-02 v0.1.0 is a draft. A future v0.2.0 might narrow catalog families after the data audit. A v1.0.0 would require approved semantics and a bounded release decision. Version labels do not create approval; status and decision records do.
Map dependencies to clauses
Maintain a table from contract clause to data schemas, domain rules, retrieval/ranking versions, model/provider versions, prompts, application controls, interface components, evaluator versions, and operating policies. When a dependency changes, replay only the evidence whose assumptions it can affect, plus system-level regression coverage.
A model update may affect explanation criteria and candidate ordering but not a deterministic purchase prohibition if tool permissions are truly separate. A catalog schema change may affect nearly every evidence clause even when the model is unchanged. This map avoids both careless under-testing and ritual full revalidation with no causal reasoning.
Preserve evidence lineage
A release record should identify the contract version, component versions, evaluation-set version, evaluator version, results, limitations, disposition, and authority decision. If production behavior later regresses, the team can reconstruct which claim was justified under which evidence.
Do not overwrite old evaluation results when a criterion changes. The old result answered the old contract. Keep it available and mark why it no longer supports the current claim.
Treat incidents as contract tests
An incident can reveal an implementation violation or a missing clause. If Patchwork falsely implies compatibility despite an existing prohibition, fix the implementation and strengthen detection. If the behavior was not prohibited because the team never considered a stale-unit state, add or revise the contract, create cases, and replay affected evidence.
The goal is not to predict every future failure. It is to make learning produce a controlled contract change rather than an isolated prompt patch.
Retire clauses deliberately
A clause can disappear because a task segment is withdrawn, a mechanism is removed, or a stronger invariant replaces it. Record the reason and ensure no interface, evaluator, control, or runbook still depends on the old ID. Silent deletion destroys the evidence trail.
Work through an authority dispute
Suppose product leadership wants Patchwork to infer a missing unit for legacy listings because abstention coverage is high. An engineer shows that most values near 16 likely mean millimeters. A catalog specialist says some older sellers used inches. A launch date is close.
The wrong resolution is a majority vote among people in the meeting. The behavior contract provides a decision path.
First, state the affected clause: PF-B06 prohibits inference when alternative units change compatibility. Second, state the evidence: sampled values, seller history, known schema periods, and segment-specific error. Third, state the limitation: prevalence is uncertain and an incorrect inference can create an unsupported fit claim. Fourth, identify alternatives: keep abstention, narrow to catalogs with explicit units, obtain an authoritative migration rule, or change the product task to similarity-only discovery. Fifth, route the domain semantic decision to the catalog domain owner and the exposure decision to the designated product/risk authority.
The Applied AI Engineer can recommend narrowing. They cannot transform statistical likelihood into domain authorization. If the domain owner cannot establish a safe inference rule, the system remains clarifying or abstaining for that segment. Deadline pressure changes cost; it does not change evidence semantics.
Now suppose the catalog owner approves a bounded inference for listings from one controlled importer and date range. The contract should not say “missing units may be inferred.” It should add the precise source segment, rule version, evidence display, test cases, correction path, monitoring, and replay trigger. Authority becomes an inspectable rule, not a verbal exception.
This method also prevents the opposite problem: risk language freezing a system without a decision. The contract forces an owner to choose among explicit scoped alternatives and records what evidence would permit expansion.
Design the fallback as a real product
Fallback is often drawn as an arrow to “manual” or “search.” That is insufficient. A fallback must preserve task continuity and an honest promise.
For Patchwork’s ordinary-search fallback, specify:
- which query and filter state transfers;
- whether user-entered measurements remain local and visible;
- which assisted labels disappear;
- how the product explains the transition;
- whether unavailable evidence is distinguished from no matching listings;
- how a user returns after correcting input;
- which events record degradation without collecting unnecessary content;
- what support sees;
- what ends the degraded state.
If assisted behavior fails after candidates appear, do not leave stale explanation text beside re-ranked ordinary results. Either freeze the evidence with a visible version or clear the assisted representation. The fallback must not combine states in a way that implies continuity it cannot prove.
A manual-review fallback needs even more detail: queue owner, competence, evidence packet, response time, capacity, permissions, allowed decisions, audit, user status, expiration, and behavior when the queue is unavailable. If those do not exist, manual review is not a fallback. Abstention or scope reduction is more honest.
Fallback evidence belongs in evaluation. Inject dependency failures, verify state transfer, and observe whether a user can continue. Reliability is part of behavior, not an infrastructure appendix.
Review contract quality before implementation
Use a structured review with participants who hold distinct evidence and authority. Product checks that clauses represent the intended task and promise. Domain owners check semantics and critical segments. Design and research check expectation, correction, burden, and fallback interaction. Platform and application engineers check enforceability and failure behavior. Privacy, security, trust, safety, legal, and risk roles review the decisions assigned to them. The Applied AI Engineer maintains traceability and resolves technical ambiguity.
The review should answer:
- Can every required behavior be observed at the combined-product boundary?
- Does each clause distinguish evidence prerequisite from model score?
- Are allowed variations broad enough to avoid exact-answer theater?
- Do prohibited effects have enforceable application controls?
- Can clarification end, or might it trap the user in a loop?
- Does abstention provide a usable next action?
- Is escalation credible at expected volume and consequence?
- Does degraded mode preserve an honest promise?
- Are critical segments explicit rather than hidden in averages?
- Does every reviewer have named competence and every approver a real decision right?
- Can each change trigger reach a replayable evidence set?
- Can a release disposition be made without inventing authority?
Invite one reviewer to argue that the product should be simpler. Invite another to construct the strongest unsafe behavior that still follows the literal words. The first exposes clauses that exist only to justify complexity. The second exposes semantic loopholes.
Contract quality also depends on readability. A machine-readable record should link to plain-language rationale and examples. A domain owner should not need to inspect code to understand the rule they are authorizing. An engineer should not need to interpret policy prose to learn which state to implement. Shared IDs connect the views.
Resolve comments through decision records rather than silent edits. If product wants more coverage and domain review wants stronger abstention, record the contested consequence, alternatives, evidence, and authority. The contract version should reveal the resolution.
Finally, review what the contract does not cover. Patchwork v0.1 excludes high-consequence categories, seller contact, and purchase. Tests should verify those boundaries, but the team should not design imaginary internal behavior for them. Non-scope prevents accidental capability growth and keeps future expansion subject to a fresh task and authority case.
A complete clause in prose
Consider PF-B13, the seller-question draft. A weak requirement says, “The AI may help users message sellers.” The complete clause is narrower.
The state begins only after a buyer explicitly requests a draft for a candidate already visible in the allowed discovery path. The draft may ask for a missing named attribute and may include values the buyer entered. It must distinguish buyer-provided from listing-provided facts, remain editable, and show that it has not been sent. It may not pressure the seller to guarantee fit, invent a part fact, access private history, choose a recipient beyond the selected listing, or call a messaging tool.
Deterministic controls remove send permissions from the companion path. Semantic cases check whether draft wording preserves provenance and avoids prohibited claims. Interaction tests check that users understand the draft state. Trust and product owners review the allowed template; any future send behavior requires a new task, privacy/policy analysis, permission model, abuse controls, and authorization.
If generation is unavailable, the degraded state offers a structured editable question template. If required evidence is absent, the state clarifies rather than drafting around it. A model or prompt change replays semantic cases. A tool-permission change blocks release until the non-contact invariant is verified.
That paragraph is longer than “help write messages” because it defines behavior at the places where implementation can otherwise invent policy. The table and JSON instance make the same decision traceable for tools; the prose preserves why the boundary exists.
Common contract failures
Quality adjectives without observables
“Accurate, relevant, safe, transparent” remains untestable.
Repair: write trigger, criteria, unacceptable behavior, evidence, state, evaluation, and disposition.
Never abstain
Product pressure treats coverage as usefulness and forces guesses.
Repair: measure coverage-risk and user burden, then narrow task or improve evidence. Do not hide uncertainty.
Confidence equals correctness
A raw score becomes a user-facing probability or release guarantee.
Repair: validate calibration for the precise outcome and segment, and keep semantic and authority conditions separate. [CLM-013]
Generic human in the loop
Review exists only as a diagram box.
Repair: name competence, evidence, timing, reversibility, capacity, fallback, and decision right.
Provider policy equals product policy
The application inherits a model specification as its contract.
Repair: treat provider behavior as one volatile component dependency; retain application clauses and replay on change. [CLM-014]
Schema equals truth
The output parses, so the team considers it correct.
Repair: distinguish syntactic, semantic, evidence, authorization, and effect validation.
Self-authorization
The implementation owner decides that test results make a risk acceptable.
Repair: attach evidence and recommendation to the named authority; block when it is absent.
Completion exercise
Write fifteen clauses for an AI-assisted task. Include at least:
- two ordinary allowed variations;
- two missing or conflicting evidence states;
- one critical segment;
- one clarification;
- one abstention;
- one escalation with a credible reviewer;
- two degraded modes;
- two prohibited effects;
- one user correction path;
- one change-triggered replay.
For every clause, supply an example, counterexample, evaluation method, release implication, and authority. Then attack the contract with a system that follows its syntax while violating its intent. Revise until the failure is observable.
The exercise fails if “human review” is unnamed, if confidence is treated as correctness, if generation is allowed to override a deterministic exclusion, or if the engineer approves their own residual risk.
Chapter decision
A behavior contract is the versioned application boundary between product intent and uncertain implementation. It specifies required, allowed, clarifying, abstaining, escalating, degraded, and prohibited states; links each clause to evidence and release consequence; and separates technical review from formal authority.
Patchwork now has PF-02 v0.1, including fifteen clauses, non-goals, an authority map, and a provider-neutral machine-readable schema. Its strongest rule is simple: the system may propose evidence-backed candidates, but it cannot create evidence, overrule known incompatibility, imply guaranteed fit, contact a seller, or purchase.
The contract deliberately does not choose a model. Chapter 4 will compare improved ordinary search, deterministic filters, hybrid retrieval and ranking, generative explanation, and bounded draft assistance against the same clauses. A mechanism earns inclusion only if it improves the contracted behavior enough to justify its new uncertainty, cost, latency, control, and change burden.