Release, Observe, and Change the Model
Turn the complete behavior dossier into an authority-bound release, observation, rollback, migration, and adaptation-referral system.
Release is an evidence boundary
Mosaic Desk now has fifteen chapters of artifacts and no permission to skip a gate. The fictional team has a task and authority contract, a selected access posture, a replayable baseline, typed proposals, bounded context, authorized retrieval, provenance-aware assembly, a segmented case set, calibrated judgment methods, and controlled experiment evidence. Chapter 15 ended with revise, not release, for a candidate instruction bundle that improved an aggregate while regressing a Gujarati citation case and a synthetic tail-duration gate.
That result does not prevent the existing qualified baseline from entering release review. It prevents the losing candidate from replacing it. The distinction matters: a release unit is an identified system configuration, not the newest model, prompt, or branch. The incoming MD-07 packet therefore carries accepted baseline evidence, rejected candidate evidence, visible regressions, and the limits of both.
This chapter closes Volume 1 by turning the complete dossier into five connected controls:
- a provider-neutral service boundary;
- a readiness packet whose gates have evidence and owners;
- a privacy-minimized observation and incident path;
- a reversible model/provider migration replay;
- an adaptation referral that can say no.
Mosaic Desk remains constructed teaching material. The companion stages no traffic, calls no provider, collects no user content, and performs no external effect. Its statuses show decision mechanics, not a production outcome.
Freeze the release unit before reviewing it
Create one release-candidate identity that links every behavior-facing dependency:
- task and authority contract;
- service-adapter and access-posture version;
- provider-exposed model identifier or open-weight checkpoint digest;
- tokenizer, chat template, messages, decoding, and schema;
- corpus, chunking, filters, reranker, assembler, and citation policy;
- case-set, rubric, taxonomy, judge, and experiment versions;
- runtime, region, capacity, retry, timeout, and fallback configuration;
- control tests, observation schema, and rollback target.
If any identity is unknown, the candidate is not one reproducible release unit. A managed alias that can change invisibly is a declared limitation. An open-weight server rebuilt with a different tokenizer or quantization is a different unit even if the checkpoint name remains the same. A changed retrieval policy is a system change even when the model is untouched.
The release packet should point back to raw evidence rather than copy scores into a presentation. A score without its case-set version, denominator, judge, segment, and failure transitions cannot support rollback or migration comparison.
Put every access posture behind one semantic seam
The application should send and receive the same semantic objects regardless of provider. Mosaic’s request carries authorized intent, trusted control, untrusted case data, authorized evidence envelopes, and deterministic state. Its response is a typed proposal with evidence references, uncertainty, and an explicit terminal state. It cannot approve warranty, contact a customer, order a part, schedule work, or mutate a record.
A managed adapter maps this seam to a provider request and records the provider-exposed model and options. The provider operates the underlying service and publishes its lifecycle surface. Mosaic still owns request mapping, authorization before context, output validation, evaluation, and the evidence supporting a change.
An open-weight adapter preserves the same application contract but needs more identity: checkpoint and tokenizer digests, chat template, runtime, kernels, quantization, hardware, batching, and scheduler state. Platform/SRE owns the serving substrate and capacity operation. The LLM engineer owns behavior-facing evidence and compatibility analysis, not cluster operation.
This separation enables comparison without pretending the paths are identical. Behavior can be replayed against one task contract. Resource evidence must retain the environment that produced it. Provider opacity and self-hosting responsibility are limitations to record, not reasons to abandon a common interface.
Assemble a readiness packet with named authorities
Release evidence spans behavior, reliability, latency/cost, privacy/security controls, owners, observability, and rollback. NIST’s voluntary profile, current evaluation practice, latency guidance, and safety guidance motivate inspecting these dimensions together, but none defines Mosaic’s organizational approval or proves a control sufficient. The checklist in this chapter is an engineering synthesis, not a certification. [CLM-046]
Build the packet as gates, not a folder of reassuring documents.
| Gate | Required evidence | Decision owner |
|---|---|---|
| behavior | frozen cases, raw transitions, segments, abstention, errors, judge limits | LLM engineering recommends; product and domain owners accept scope |
| authorization and threat controls | tests at retrieval, assembly, output, and effect boundaries; residuals | security and system owners |
| privacy | signal purpose, fields, access, retention, deletion, sampling | privacy authority |
| reliability and resources | failures, retries, timeouts, tails, load assumptions, capacity, fallback | platform/SRE |
| domain | criteria, consequential cases, escalation, reviewer competence | domain authority |
| release | linked gate dispositions, cohort, stop triggers, recovery evidence | designated release authority |
A gate can be evidence-attached, authority-review-required, blocked, or approved-by-named-authority. The engineer cannot translate a missing signature into implied approval. Deadline pressure does not change evidence state.
The deterministic companion intentionally leaves privacy, security, platform, domain, and release gates awaiting their authorities. Its result is hold-for-authority-review. That is a complete technical handoff, not a failed engineering task.

CLM-046 and CLM-048; it reports no live traffic or quality.Long description
A colorful three-dimensional service path carries one case envelope through an interface gate, authorized retrieval shelves, a model chamber, schema and citation validation, and a human decision desk. Small side gauges receive only version, state, segment, reason-code, and duration buckets. The proposal path stops before any external-effect system.
Observe decisions without collecting the conversation by default
Start signal design from a decision. Ask what evidence would cause the team to pause, contain, replay, roll back, or investigate. Do not begin by storing every prompt and response.
A minimal event can contain a one-way trace reference, system version, coarse authorized segment, terminal state, reason code, citation-validation state, duration bucket, and time bucket. It need not contain the user’s words, retrieved passages, generated prose, direct identity, or unrestricted metadata.
For every signal, record:
- purpose and exact decision it supports;
- allowed fields and prohibited payloads;
- sensitivity and access boundary;
- owner and first response;
- retention and deletion rule approved by privacy authority;
- segment denominator and sampling method;
- threshold or comparison rule;
- blind spots and escalation route.
Hashes and redaction are not magic anonymity. A stable trace hash can still be linkable. A rare segment can identify a person indirectly. Privacy owners decide necessity, access, retention, and deletion. Security owners decide whether a signal or trace creates a threat surface. Domain owners decide when a sampled case needs qualified review.
Raw traces may be justified for a bounded investigation, but that is a separate access-controlled collection with purpose, minimization, expiry, and approval. It is not the default observability strategy.
Measure system states and resource tails
The observation system must preserve the behavior states defined in Chapter 2 and implemented in Chapter 7. Report success, abstain, degraded, fail-closed, and escalate independently. A lower success rate can be desirable if unsupported answers become correct abstentions. A low error rate can be dangerous if unauthorized evidence silently passes.
Track retrieval eligibility and exclusion reasons, evidence states, citation validation, schema outcomes, repair attempts, terminal states, and human escalation. These are component signals. They do not prove source truth or final domain correctness.
For resources, separate queue, retrieval, model, validation, and end-to-end duration where the path exposes them. Preserve tails and failures. Record input/output units, retries, throttling, cache state, concurrency, and estimated or observed cost under a versioned environment. Current provider latency advice is useful implementation guidance, not a Mosaic capacity result or a portable promise.
Define budgets with behavior under breach:
- before model execution, reject or route when required evidence cannot fit;
- on a deadline, stop, abstain, degrade to an approved deterministic path, or route to review;
- after an unknown provider outcome, reconcile before retrying any downstream effect;
- when a protected segment or authorization gate regresses, pause the affected path regardless of aggregate health.
Platform/SRE owns service-level objectives, capacity, alert operation, and infrastructure recovery. LLM engineering supplies behavior-layer signals and diagnostic evidence. The two responsibilities meet at the release packet; neither disappears into the other.
Progress through bounded release states
Use an explicit state machine:
- Offline: replay frozen cases, failure injections, and deterministic controls. No user sees output.
- Internal: authorized reviewers inspect proposals and escalation behavior. No external effect follows.
- Shadow: the candidate receives authorized mirrored inputs under approved privacy policy, but its output is not shown or acted upon.
- Bounded cohort: only approved segments and routes receive proposals, with hard stop triggers and a qualified fallback.
- Expanded scope: require new evidence and authority for each material expansion. Do not treat time in service as proof.
The next state is not automatic. Each transition names entry evidence, decision owner, cohort scope, monitoring, duration or sample condition, stop triggers, and recovery path. Shadow operation is not harmless if it collects disallowed data or consumes unsafe capacity. A small cohort is not safe if it contains a high-consequence segment with an unresolved gate.
Rollback must be executable before progression. Keep the previous qualified adapter and its compatible dependencies available where feasible. If exact rollback is impossible because a provider endpoint is gone, define a tested fallback such as a different qualified adapter, deterministic information-only behavior, correct abstention, or a human queue. Never label a plan rollback when it cannot restore a coherent contract.
Detect, contain, then diagnose the responsible layer
An alert is not a diagnosis. First contain the possible consequence: pause the candidate, exclude a segment, fail closed, restore the qualified adapter, or route to review. Then compare evidence across layers.
Inspect in this order:
- task/authority contract and request classification;
- input validation and language/segment routing;
- source eligibility and permission;
- query, candidate retrieval, filters, and reranking;
- context assembly, source state, ordering, and truncation;
- message/template, model, and decoding identity;
- schema, provenance, citation, and deterministic controls;
- judge/rater identity and disagreement;
- runtime, caching, retry, throttling, and capacity state;
- interface rendering and downstream handling.
Use interventions to separate causes. Replay the same case with the frozen corpus, then the candidate corpus; the current adapter, then candidate adapter; the baseline judge, then a reviewed deterministic assertion. Do not blame the model because its output is the most visible artifact.
LLME-CASE-014 supplies a risk pattern: untrusted content may carry competing instructions, so model output stays untrusted and authorization remains outside the model. It does not prove a prevention rate or make prompt text a security boundary. Current provider safety guidance likewise cannot approve Mosaic’s controls. Security authority owns threat decisions and residual risk.
Incident command remains with the designated operational owner. The LLM engineer can reproduce behavior, locate a failing layer, recommend containment, and update evaluation. They do not become incident commander, privacy officer, domain authority, or platform operator by writing the model integration.
Treat provider lifecycle as a normal change trigger
LLME-CASE-013 demonstrates only that official managed-service deprecation schedules exist and change. It reports no Mosaic outcome and endorses no provider. Monitor official notices, preserve the dated snapshot used for planning, and treat the proposed replacement as an unqualified dependency.
Model/provider migration requires replay of frozen task evidence and recording changed limitations. Current provider evaluation and lifecycle documentation support requalification and change tracking, but a listed replacement is not evidence of behavioral equivalence. Deprecation is one trigger among model revision, endpoint change, policy change, template change, runtime rebuild, quantization, hardware change, or observed regression. [CLM-047]
Build a migration compatibility matrix across:
- semantic request and response contract;
- supported schema and terminal states;
- tokenizer/template and context behavior;
- retrieval/context integration and citation handling;
- frozen case and protected-segment results;
- judge compatibility and raw disagreement;
- latency, failure, retry, rate-limit, and cost evidence;
- privacy, security, region, and data-handling constraints;
- lifecycle, rollback, and fallback limits.
Run the candidate in replay and shadow before any cohort. Keep current and candidate identities separate, preserve pass-to-fail transitions, and record every limitation that changes. A candidate can improve one slice while becoming unacceptable overall.

CLM-047 and the Volume 2 handoff; it is not migration-success evidence.Long description
Two realistic three-dimensional adapter modules labeled current and candidate connect to one frozen case vault and judgment console. The candidate travels through replay and shadow checkpoints. A changed-limits ledger records Gujarati citation and tail-duration regressions, while a bright rollback rail returns traffic eligibility to the current adapter. No path advances directly to production.
Failure injection: two regressions are not one model problem
The synthetic migration fixture injects a provider retirement and a candidate adapter that appears interface-compatible. Under three constructed replay cases, English formatting changes from fail to pass. A protected Gujarati citation case changes from pass to fail. Gujarati-English conflict escalation remains correct. Candidate p99 teaching duration rises from 90 to 135 synthetic units against a gate of 110.
The disposition is hold-and-keep-current-fallback. The numbers demonstrate decision logic only. They are not live latency, language-quality, or provider evidence.
A second injection allows a revoked source into the rerankable set. This is a retrieval-authorization regression, not proof that the candidate model caused unauthorized evidence. Contain it by failing closed and revoking the affected corpus/policy version. Diagnose the authorization path separately from the provider adapter.
The incident record therefore carries two layer labels, owners, containment paths, and repair queues. Collapsing them into “the new model is worse” would send engineering effort to the wrong surface and hide a security-relevant defect.
Reject adaptation until the residual earns investigation
Adaptation is justified only after residual failures survive simpler task, context, retrieval, prompt, schema, and evaluation repairs. Evaluation guidance and the original RAG evidence support iterative system diagnosis, but they do not define a universal intervention ladder or prove that changing weights will fix a failure. Intervention order depends on the observed mechanism, evidence, access, and constraints. [CLM-048]
An adaptation referral must answer five questions:
- Stable: Does the same bounded failure recur across identified runs, cases, and environments rather than one migration sample?
- Valuable: Does a product and domain owner confirm that fixing the residual matters for an approved scope?
- Data-supported: Is there authorized, representative training and evaluation evidence that expresses the desired distinction?
- Adaptation-sensitive: Is there a plausible reason weight change could address the residual without a smaller interface or evidence repair?
- Protected: Are retention, control, multilingual, abstention, and resource regressions explicit and release-blocking?
If any answer is no or unknown, the referral is not-justified, repair-first, or gather-evidence. Fine-tuning is not a reward for exhausting patience with a prompt.
Mosaic’s provider-replacement Gujarati citation regression fails the gate. It is not stable across runs, product/domain value is undecided, training data are not established, adaptation sensitivity is unsupported, and simpler adapter, template, retrieval-permission, context, citation-validation, and replacement-selection repairs remain. Volume 2 receives this negative disposition and must audit it. It receives no implied approval to train.
Build the MD-08 handoff
The final Volume 1 dossier contains:
- service-adapter contract and managed/open-weight responsibility matrix;
- complete release-candidate dependency identity;
- gate/evidence/owner readiness packet;
- behavior, resource, control, and privacy-minimized signal catalog;
- offline, internal, shadow, and bounded-cohort state plan;
- stop triggers, fallback, rollback, and recovery limits;
- incident roles, layered trace, containment, and diagnostic evidence;
- provider notice snapshot, compatibility matrix, replay, shadow, and changed limitations;
- adaptation gate with stable, valuable, data-supported, adaptation-sensitive, and protected-case fields;
- explicit unknowns, residuals, dispositions, and authority handoffs.
This completes MD-08. The release packet is ready for named-authority review; it is not self-approved. The migration candidate is held. The retrieval authorization defect is separately contained. The adaptation referral is not justified. Every status is useful because it preserves what evidence can and cannot support.
Practice: defend release and change without borrowing authority
Use the Chapter 16 companion to produce a review packet:
- reconstruct the release-candidate identity from Chapters 2, 5, and 9 through 15;
- map each gate to evidence, limitation, owner, and unresolved decision;
- inspect every signal for purpose, minimization, access, retention, decision, and blind spot;
- explain why the stage order cannot be skipped;
- rehearse a protected-segment rollback or qualified fallback;
- diagnose the provider-adapter and retrieval-permission injections independently;
- defend the migration hold using raw transitions and resource limits;
- complete the five-question adaptation referral;
- state what Volume 2 may investigate and what it may not infer;
- name every product, domain, privacy, security, platform/SRE, incident, and release handoff.
Pass when another reviewer can reproduce the configuration, find the evidence behind each gate, identify the first safe containment for both injected failures, execute a coherent rollback or fallback, and explain why neither the release packet nor migration replay grants production or adaptation authority.
Fail if raw content is collected without purpose and approval, aggregate health hides a protected regression, the candidate silently replaces the baseline, rollback restores an incompatible dependency bundle, two failure layers are merged, or adaptation is recommended from a novel but unstable residual.
Volume 2 begins only after root acceptance of this Volume 1 dossier. Its first chapter must reconstruct MD-08, preserve the negative adaptation disposition, and test whether a valuable residual remains after simpler repairs. It may change the conclusion only with new traceable evidence.