Release, Diagnose, and Evolve Adapted Models
Close the series with an authority-preserving release, incident, rollback, requalification, and durable-learning loop across the complete model-system dossier.
A complete dossier can still say hold
MD-16 consumes every Mosaic artifact from the language-task contract through the simulated serving frontier. The evidence chain is complete enough to reconstruct, but the open-weight package is not executable: the adapter was never trained and the template remains incompatible. The readiness disposition is therefore hold. This is a successful evidence process without a fabricated release.
Adapted-model release needs target, retention, control, runtime, owner, observability, cohort, and rollback evidence. [CLM-049] The packet links exact package identity, data/training lineage, compatibility, behavior/resource frontier, raw protected cases, limitations, privacy review, named decision owners, staged cohort, stop triggers, prior recoverable package, and reconciliation plan.
Offline pass is necessary but not sufficient. The state machine is offline-qualified, shadow, internal, restricted cohort, wider cohort, release, hold, rollback, or deprecate. Each transition names entry evidence, observation window, stop rules, decision owner, and fallback. No deadline changes the evidence label.
Because the package is blocked, Mosaic performs a tabletop only. No user traffic, model call, deployment, or external effect occurs.

Long description
A realistic control-room gate accepts separate evidence keys from behavior, runtime, privacy, owners, cohort, and rollback stations. A hold lamp remains lit because the package key is incomplete.
Observe decisions without collecting the product
Monitoring begins with decisions: what signal triggers contain, diagnose, rollback, or review? Record request class, package and configuration identity, language/slice, schema state, evidence/citation status, abstention, fallback, latency stage, resource state, error code, and cohort. Prefer counts, bounded categories, hashes, and sampled approved fixtures over raw prompts or outputs.
Each signal names purpose, sensitivity, access, retention, aggregation, threshold, owner, action, and blind spot. Rare or high-consequence failures need protected review paths; unrestricted logging is prohibited. Privacy authority decides permissible collection. Monitoring distributions can drift and do not prove causality.
The tabletop failure begins after a hypothetical partial cohort: a tokenizer/runtime update lowers synthetic memory but changes truncation, degrading long Gujarati behavior. An adapter/template mismatch is a competing explanation. The first action is contain: stop ramp, route to the prior qualified fallback, preserve privacy-minimized identities and traces, and open an incident timeline. Do not retrain or retune before isolating the layer.
Diagnose by layer
Incident diagnosis must distinguish data, checkpoint or adapter, tokenizer or template, evaluator, runtime, and infrastructure changes. [CLM-050] Some failures cross layers, so test hypotheses rather than assigning blame.
The diagnostic ladder asks for the next discriminating evidence:
- input/data: distribution, language, authorization, freshness, malformed state;
- tokenizer/template: revision, serialization, truncation, special tokens, padding;
- retrieval/context: source snapshot, permissions, conflicts, assembly, missing evidence;
- base/adapter/merge: exact digests, binding, checkpoint, merge state;
- decoding: sampling, stops, schema constraints, fallback;
- runtime/kernel: version, precision, quantization, kernel path, cache, scheduler;
- evaluator: prompt/model/version, calibration, disagreement, slice coverage;
- product integration/infrastructure: routing, timeout, retries, state, capacity.
Replay the same frozen cases while changing one suspected surface. The tabletop compares old and new tokenizer serialization under the same package plan and finds earlier truncation of Gujarati citation evidence. Restoring the pinned tokenizer removes the deterministic failure. This supports a tokenizer/runtime cause in the fixture; it is not a production diagnosis.

Long description
A colorful three-dimensional incident loop surrounds a layered diagnostic tower. The path can return to rollback or proceed through authority decision to durable learning.
Correct, verify, and preserve rollback
A correction is incomplete until evidence, tests, controls, manifests, limitations, and monitoring are updated and replayed. [CLM-051] Pinning the prior tokenizer is containment or rollback, not necessarily the long-term correction. A forward fix creates a new complete identity and replays compatibility, target, retention, language, control, workload, degradation, and recovery evidence.
The rollback record binds the previous package, trigger, routing change, state reconciliation, verification cases, owners, and outcome state. Unknown external effects require reconciliation before retry. Mosaic’s tabletop performs no effect and records rollback as simulated.
Deprecation needs last-supported identity, replacement or fallback, migration evidence, owner, archive/retention, and unresolved users. Provider deprecation schedules are volatile dependency inputs. A managed fallback is versioned and requalified under the same external behavior contract, without invented internals.
Keep authority distributed
LLM engineering owns model-system evidence, layered diagnosis coordination, and replay recommendations. Platform/SRE owns capacity, routing, reliability, and operational recovery. Incident command owns incident decisions and communications. Product/domain/language authorities own behavior consequence; privacy owns observation; security owns security response; legal owns legal interpretation; release authority issues ramp, hold, rollback, or deprecation.
The engineer cannot self-approve because tests pass, and a monitoring dashboard cannot accept residual risk. Disagreement remains in the record with evidence, owner, deadline, and escalation path.
Close MD-16 and the series
The deterministic release artifact links MD-01 through MD-16, issues a hold for the blocked package, runs a tabletop tokenizer/runtime incident, demonstrates layered diagnosis and rollback evidence, and updates tests, card, limitations, monitoring, and change ledger. No Mosaic outcome is claimed.
Portfolio learning separates durable reusable controls from local choices. Reusable candidates include manifest identity, compatibility gates, protected-slice replay, privacy-minimized signal contracts, and layered diagnosis. Local data, thresholds, language judgments, capacity, and release decisions do not become universal standards from one fictional case.
The closing case set remains bounded: LLME-CASE-008 supplies staged-recipe lineage, LLME-CASE-009 adapter compatibility, LLME-CASE-011 serving workload reasoning, LLME-CASE-012 compression requalification, LLME-CASE-013 lifecycle change, and LLME-CASE-014 untrusted-content risk. None supplies a Mosaic release, incident, safety, or capacity outcome.
Volume 1 taught how to define, build, evaluate, release, and change observable LLM behavior. Volume 2 began only after the MD-08 adaptation referral, ruled out upstream defects, built data and experiment evidence, rejected unjustified optimization, preserved package truth, and closed with a reversible hold. The final lesson is not “always tune.” It is to make every change earn its place in a reconstructable behavior system.
Practice
Run the tabletop from detection through hold or rollback. Pass when the causal layer is supported by controlled replay, privacy-minimized signals remain adequate, rollback is reconstructable, formal owners sign their decisions, uncertainty stays visible, and the entire MD-01 to MD-16 chain can be rebuilt.