NEWProduction Web Themes & Turnkey ArchitecturesGet Lifetime Pass ($199) →
KNKomal Nakrani
Get All Access
ThemesDocsAll-Access PassGet All Access ($199)
Book overview
13/LLM Adaptation and Runtime

Decide on Distillation or Continued Pretraining

Choose advanced adaptation only when its mechanism, transfer source, authority, and disconfirming evidence match a distinct residual gap.

Escalation needs a distinct thesis

MD-13 opens with the accepted preference rejection. The remaining ledger names bounded shorthand coverage, unresolved template compatibility, and insufficient independent language evidence. None automatically implies distillation or continued pretraining.

Begin with alternatives: keep no-tune, repair prompt/context/retrieval, use bounded SFT or PEFT, replace the base or managed comparator, or stop. Ask whether the gap is stable behavior, missing distribution, capacity, or changing knowledge; what smaller path failed; what teacher or corpus carries the signal; and what held-out evidence could disconfirm the method.

Residual diagnosis routes separately to repair, SFT or PEFT, distillation, continued pretraining, replacement, no project, or rejection.
V2-F13.1 - Methods answer different failure hypotheses. Essential labels: residual, repair, SFT/PEFT, distill, continue, replace, reject. Branch shapes supplement color. Supports CLM-037 to CLM-039; not a ranking.
Long description

A colorful three-dimensional junction routes an inspected residual toward separate method bays, each with an evidence gate and rejection buffer.

Distillation transfers teacher behavior and defects

Distillation transfers teacher behavior into a student under an explicit objective and data process; it is not generic compression magic. [CLM-037] Teacher distributions, generated targets, or demonstrations remain bounded by teacher quality, student capacity, data, decoding, objective, and evaluation.

A teacher ledger pins model or service version, prompt, generation settings, inputs, filtering, rights, cost, unavailable internals, and limits. Targets that make grounded claims need authorized evidence links. Mosaic’s failure injection produces fluent policy targets containing unsupported details. They are rejected before a pilot; grammar filtering would preserve the transferred defect.

Distillation is plausible for a stable bounded behavior when a smaller student must meet a resource goal and independent suites can detect lost capability. It is mismatched for changing knowledge, broken templates, or missing authorization. A bounded pilot would cap calls, updates, data, time, storage, and evaluation, then compare teacher, student baseline, distilled candidate, and the existing best path. Teacher agreement is not acceptance.

Continued pretraining changes a distribution

Continued pretraining targets a domain or task distribution through additional unsupervised language modeling and carries contamination and retention costs. [CLM-038] It may fit a demonstrated language or domain representation gap with an authorized corpus. It does not fit frequently changing policy that belongs in retrieval.

A corpus ledger records sources, permission facts for legal interpretation, dates, language/domain mixture, transformations, deduplication, restricted-data review, contamination, split relationships, hashes, retention, and omissions. Dolma (LLME-CASE-006) and domain-adaptive pretraining research provide bounded process evidence, not Mosaic authorization or forecasts.

The failure injection proposes memorizing volatile policy. Mosaic rejects it: weights obscure freshness and correction while retrieval already owns current authority. Independent target, retention, language, abstention, citation, control, and no-change suites remain required because corpus loss cannot select product behavior.

Teacher and corpus paths cross provenance, rights, contamination, and independent evaluation barriers before reaching a student or base model.
V2-F13.2 - Transfer sources remain bounded by lineage and tests. Essential labels: teacher, corpus, provenance, rights, contamination, evaluate. Supports CLM-037 and CLM-038.
Long description

A realistic lab routes teacher cards and corpus containers through sealed provenance, rights, contamination, and evaluation gates.

Reject fashion with common evidence

Advanced adaptation should be rejected when its hypothesis, data, resource, or evaluation burden is not distinct from a smaller intervention. [CLM-039] Scaling and staged-recipe research describes studied regimes; it cannot forecast Mosaic fitness from parameters, tokens, or method names.

Compare residual mechanism, transfer authority, target/protected evidence, resources, compatibility, rollback, change cadence, uncertainty, and owners on one sheet. Research owns novel methods; legal/privacy/domain authorities approve transfer sources; platform owns compute; LLM engineering owns method fit and behavior evidence.

The deterministic companion rejects continued pretraining because policy is volatile and rejects distillation because teacher targets lack evidence and add no distinct value. No pilot runs. Chapter 14 receives the simulated peft-r8-attn lineage with an explicit non-executable state.

Design a pilot that can stop

If a future residual clears the method gate, the pilot record freezes the question before computation. For distillation it names the teacher and student identities, target-generation route, temperature or decoding policy, evidence validator, rejected-target archive, student objective, and teacher-call budget. For continued pretraining it names corpus identity, sampling mixture, tokenizer, sequence policy, objective, update budget, checkpoints, and contamination scan. Both routes inherit the exact baseline, protected suites, and package compatibility tuple.

Each pilot has automatic stops: unauthorized source access, identity or template mismatch, contaminated holdout, non-finite optimization, resource breach, invalid checkpoint, forbidden retention transition, target futility, or authority withdrawal. A stop is evidence about the thesis; it is not a reason to widen scope mid-run.

Resource comparison includes data acquisition and review, teacher generation, training memory and time, evaluation, storage, inference consequences, human adjudication, and maintenance. Distillation that reduces inference cost may add recurring teacher and refresh costs. Continued pretraining may create a larger artifact and more expensive future adaptation. Report the whole lifecycle rather than only training parameters.

Preserve the no-project option

The decision sheet records what happens if Mosaic does nothing. The current no-tune or simulated PEFT path may remain preferable when evidence is insufficient, the gap is low value, or the authority and maintenance burden outweigh likely benefit. “No project” is not failure; it is a controlled disposition with revisit triggers.

Revisit only when a stable residual appears across fresh held-out cases, a suitable authorized teacher or corpus exists, compatibility is repaired, and a bounded pilot can distinguish the method from simpler alternatives. This prevents repeated exploration from quietly becoming an indefinite training program.

The Tulu 3 case (LLME-CASE-008) demonstrates why stage-level artifacts and comparisons matter, but its recipe does not select Mosaic’s stages. The corpus case demonstrates process and contamination concerns, not blanket permission. Every imported lesson remains attached to its model, data, task, and measurement limits.

Practice

Classify a latency-constrained stable behavior, changing policy, and language-distribution deficit. Pass only when mechanism, authority, contamination, retention, resources, and disconfirmation justify pilot or rejection.