NEWProduction Web Themes & Turnkey ArchitecturesGet Lifetime Pass ($199) →
KNKomal Nakrani
Get All Access
ThemesDocsAll-Access PassGet All Access ($199)
Book overview
08/LLM Adaptation and Runtime

Preserve Retention and Control Cases

Freeze target, retention, multilingual, abstention, safety, and no-change lanes so adaptation cannot hide forbidden regressions.

Improvement needs a protected outside

The instruction and preference datasets define candidate learning signals. They do not define acceptable collateral change. Chapter 8 completes MD-11 by freezing cases that adaptation is not allowed to silently damage.

Adaptation promotion requires target improvement plus retention, segment, and control regression evidence. [CLM-022] The protected suite is finite and incomplete. Its purpose is to make known obligations and forbidden transitions release-blocking, not to certify general capability or safety.

Mosaic retains the accepted no-adaptation baseline, unchanged retrieval/context/message/schema configuration, exact evaluator identities, and raw case outputs. The open-weight candidate may change only inside the declared experiment. The managed baseline is a semantic comparator, not a source of training targets.

Build the dependency map

For every target clause, ask what neighboring behavior could regress. A shorthand mapping change can affect near-miss codes, multilingual formatting, citation placement, abstention, conflict escalation, schema validity, and authority boundaries. Adapter changes may also alter unrelated summarization or instruction behavior.

Create five lanes:

  • target: the exact shorthand residual;
  • retention: previously qualified citation, structured output, and neutral instruction cases;
  • segments: Gujarati, English, Gujarati-English, appliance families, evidence states, and consequences;
  • controls: injection attempts, unauthorized sources, prohibited effects, schema failures, and deterministic gates;
  • no-change: frozen baseline, upstream repair, retrieval repair, and prompt/context comparator results.

Each case carries provenance, family, immutable input/expected-state identity, evaluator, threshold or transition rule, owner, consequence, and limitation. Training and development procedures cannot read protected targets.

Distinguish retention from general benchmark breadth

Retention cases protect behavior already required by Mosaic’s contract or needed to interpret the experiment. They are not a random benchmark bundle. Begin with Volume 1 qualified cases and every protected criterion frozen in MD-09, then add only cases whose dependency on the proposed intervention is plausible and whose consequence is named.

General-capability probes can reveal surprising changes, but a public benchmark score is not automatically release-blocking. State why each probe belongs, whether its data may overlap the base model, and who owns the consequence. Do not let a large public suite drown the few high-consequence local cases.

Keep retention targets inaccessible to training-data authors. Aggregate slice definitions may guide coverage, but exact prompts, evidence, and expected outputs stay in the protected vault. Near-duplicate and source-family checks must include every generated or translated derivative.

A target-improvement core is surrounded by retention, multilingual, abstention, and control shields with named owners and stop gates.
V2-F08.1 - Target gain is bounded by protected behavior. Essential labels: target, retention, multilingual, abstention, control. Shield shapes supplement color. The figure supports CLM-022 through CLM-024; it is not safety certification.
Long description

A vivid realistic three-dimensional target core sits inside layered shields labeled retention, multilingual, abstention, and control. Each shield connects to a named owner and a stop gate. A candidate arrow cannot reach promotion by passing through a damaged shield.

Freeze thresholds and transitions before candidates

A protected rule can be numeric, categorical, or case-specific. Examples include no pass-to-fail authorization transition, no successful answer replacing a correct abstention, no protected-language citation regression, no schema/prohibited-effect breach, and a preregistered tolerance for a noisy low-consequence measure.

The owner of the consequence approves the rule. Security owns unacceptable threat/control regressions; domain/product owners own consequential task/segment criteria; privacy owns data-handling constraints; platform owns resource gates; release authority decides promotion. LLM engineering executes and reports.

Do not tune thresholds after seeing a checkpoint. A changed rule creates a new experiment, rationale, and authority decision.

Represent gates in a matrix:

Lane Example rule Owner Candidate consequence
target preregistered mapping improvement product/domain necessary, never sufficient
retention no citation pass-to-fail domain reject or diagnose
segment no Gujarati protected regression language/domain reject or narrow scope with authority
control no authorization or effect breach security/system immediate stop
no-change candidate must beat smaller repair experiment owner reject unnecessary adaptation

The actual thresholds remain fictional and uncalibrated. The durable structure is that each rule has evidence, an owner, and a predetermined disposition.

Protect multilingual and domain slices

Language/domain gains can coexist with regressions outside the target mixture. [CLM-023] Aya and continued-pretraining studies provide bounded evidence that language and domain adaptation are distribution-sensitive. They do not predict Mosaic transfer.

Keep Gujarati citation and conflict cases even if the target is mixed-language shorthand. Retain English neutral tasks, other appliance families, missing/stale evidence, and code-switched inputs outside the dominant training families. Report raw transitions and family counts; an aggregate cannot cancel a protected failure.

LLME-CASE-007 supports explicit language stratification, not universal adequacy. LLME-CASE-008 supports stage-level evaluation, not a required suite.

Preserve security and authority controls

Safety/control cases and accountable owners must be preserved through adaptation rather than assumed from the base model. [CLM-024] Include instruction-bearing untrusted content, unauthorized-source requests, prohibited warranty/effect requests, schema-valid but unauthorized proposals, and conflict escalation.

LLME-CASE-014, NIST guidance, and provider safety practices identify evolving trust-boundary and mitigation concerns. They do not prove exploitability, prevention rates, sufficient controls, or acceptance. Prompt text is not an authorization boundary. Deterministic retrieval, proposal validation, and effect gates remain outside the model.

Failure injection: style wins, controls lose

A synthetic checkpoint makes answers shorter and more polished on the target slice. It also changes one correct abstention to an unsupported success, regresses Gujarati citation fidelity, and follows an instruction embedded in untrusted context.

The disposition is reject-forbidden-regressions, regardless of target-style preference. The fixture is not a trained model result; it demonstrates the gate. Diagnose each transition by layer and retain the raw record. Do not average three hard failures against a style win.

Triage begins after containment of any possible consequence. Replay the same case with the frozen baseline, candidate weights under the frozen system, and no-change controls. Inspect tokenizer/template identity, retrieval eligibility, context, schema, evaluator, and runtime before attributing the transition to adapted weights. If a control dependency changed, the experiment is confounded.

Classify the outcome as candidate defect, evaluator defect, data/split leakage, system/config change, or unresolved. A corrected evaluator does not erase the original disagreement; it creates new versioned evidence. A candidate can be rejected while diagnosis continues.

Across synthetic checkpoints, a target curve rises while retention and control curves fall through blocking thresholds.
V2-F08.2 - A rising target cannot hide protected decline. Essential labels: target, retention, control, block. Line styles and threshold gates supplement color. The figure supports the forbidden-regression decision; values are illustrative only.
Long description

A three-dimensional checkpoint path carries three distinct curves. The target curve rises, but retention and control curves cross bright block planes. The candidate is routed to a rejection bay while the frozen baseline remains available.

Keep controls unchanged

Freeze corpus, retrieval filters, context assembly, messages, decoding, schema, judge, runtime, and case identity. If one must change, open a new isolated comparison. Otherwise a candidate could appear to retain behavior because retrieval improved or appear to regress because an evaluator changed.

The no-change lane includes the frozen model plus upstream language-code and retrieval repairs already selected in MD-09. This distinguishes weight effects from smaller system interventions and preserves the right to select no adaptation.

Run all candidates against the same snapshot and record pass-to-fail, fail-to-pass, abstention transitions, schema/control transitions, and raw outputs. Report by lane, language, family, evidence state, and consequence. Repeated stochastic trials preserve the trial identity and uncertainty; a lucky pass cannot overwrite a failure.

Suite maintenance is controlled change. Add a newly discovered threat or segment through a new version with provenance and owner. Keep the old suite for historical interpretation where authorized. Never backfill a new hard case into a completed experiment and claim it was preregistered.

Complete MD-11

mosaic-retention-control-suite.json freezes target, retention, segment, control, and no-change lanes with owner/threshold matrices. It contains the synthetic forbidden-regression checkpoint solely as a gate test. Status: md11-complete-protected-suite-release-blocking.

This does not clear the Chapter 2 template blocker or authorize training. It makes Chapters 9-13 accountable to the same protected evidence. Preference labels cannot override hard controls; a finite suite cannot prove absence of forgetting or unsafe behavior.

The managed and open-weight paths share outcome clauses, case inputs where the interface permits, and terminal-state judgments. Weight-specific diagnostics, tokenizer/template visibility, and runtime state remain separate. Provider opacity is a limitation, not permission to weaken the outcome gate.

Practice

Build the matrix, preserve hashes/evaluators/config, name owners and forbidden transitions, run the synthetic target-style checkpoint, reject it, and state suite blind spots.

Pass when every case has provenance, owner, block rule, unchanged control identity, raw transition, and limitation. Fail when target-only selection wins, retrieval changes during comparison, or the suite is called certification.