Budget Context Deliberately
Allocate finite context among control, authorized evidence, history, and output while making every omission and failure state visible.
A large window is not a context design
MD-03 gives Mosaic Desk a complete typed baseline. The message contract separates trusted control from untrusted data. The output contract validates structure, semantics, provenance, and authority, then routes every candidate to an explicit state.
The next synthetic case contains a long technician thread, repeated signatures, quoted replies, two service bulletins, warranty terms, a parts record, and a request for a concise proposal. Everything can be serialized into the candidate’s advertised input capacity.
That fact answers only one question: the representation may fit under a nominal limit.
It does not establish that the current bulletin will influence the output, that a conflicting old bulletin will be recognized as stale, that quoted instruction-like text will remain untrusted, that the application reserved enough output capacity, or that latency and cost fit the workload. A context window is capacity. A context contract is a selection, ordering, trust, omission, and failure policy.
Chapter 3 introduced token and template accounting. Chapter 8 turns that accounting into MD-04: a measured budget with named source classes, mandatory and optional allocations, stress cases, and traceable omissions.
Inventory sources before counting tokens
Start with semantic objects, not one concatenated string.
For each possible context source, record:
- source class and record identifier;
- owner and authority status;
- tenant/case scope and permission decision;
- freshness and supersession metadata;
- trust class from Chapter 6;
- required claim or behavior clause;
- minimum useful evidence unit;
- exact rendered token count for the selected adapter;
- mandatory, conditional, optional, or prohibited allocation;
- compression rule and what must survive it;
- omission behavior and trace field.
Mosaic’s inventory includes application control, authorized user intent, deterministic case scope, current service bulletin, relevant technician observations, conflict/supersession metadata, warranty material, parts facts, historical messages, and output reserve. A revoked or other-tenant record belongs in the prohibited class even when relevant.
Domain owners decide which sources are authoritative for each claim. Privacy and security owners decide permitted use. The LLM engineer implements the context interface and measures its behavior. A token score cannot adjudicate authority.
Write the budget equation
For a selected model/access/template identity, define:
usable input = nominal capacity
- reserved output
- template and role overhead
- fixed control and deterministic state
- safety margin
Then allocate usable input among mandatory evidence, conditional evidence, and history. Measure the exact rendered representation with the actual tokenizer when observable. For a managed path that does not expose tokenization, use provider-returned usage or documented counting where available and preserve uncertainty. Do not substitute another tokenizer as fact.
Reserve output before packing input. A request that consumes the full nominal window can truncate the response, fail at the provider boundary, or force an unplanned output cap.
The synthetic companion uses abstract units rather than claiming a real candidate’s token count. Its ledger proves arithmetic and policy invariants. A future authorized run must replace those units with actual measured behavior.

Long description
A colorful transparent case has a fixed outer boundary. Locked control blocks occupy a protected section, evidence cards with source tabs fill another, compressed history tiles occupy a conditional section, and an empty compartment labeled output remains reserved. Overflow items move to a visible omitted tray rather than vanishing.
Mandatory is about consequence, not recency alone
The packing policy should prioritize by task need, authority, and consequence.
For a safety-sensitive Mosaic case, current stop/escalation conditions and their source identity may be mandatory. The latest technician message may be relevant but cannot displace that evidence merely because it is newer. Repeated signatures and quoted history may be optional. An old bulletin that conflicts with a current record may need a compact conflict marker rather than silent deletion.
Use rules such as:
- never drop trusted application control or deterministic authorization state;
- never include unauthorized content;
- retain identity, owner, date, and supersession status with evidence;
- preserve the evidence required by each proposed claim;
- deduplicate exact or known quoted repetitions before compressing meaning;
- reserve space for conflicting evidence when the conflict changes behavior;
- record every omitted source and reason;
- enter degraded, abstain, or escalate when required evidence cannot fit.
No universal source order follows from these rules. Test ordering on the selected candidate and tasks.
Capacity does not guarantee reliable use
Long-context research has shown position-sensitive behavior in tested question-answering and retrieval tasks. LLME-CASE-002 records that relevant information was often used less effectively when placed in the middle for the studied models and setups. The case is a failure-study boundary, not a claim about every current model.
Context-window capacity does not guarantee equal or reliable use of information at every position. The magnitude and shape of position effects vary with model, task, formatting, distractors, and current implementation. Mosaic must rerun position tests against its candidate and proposal criteria. [CLM-022]
Design a position experiment:
- Hold the semantic case and configuration fixed.
- Place the decisive authorized evidence near the beginning, middle, and end.
- Preserve source identity and comparable surrounding length.
- Add a distractor condition with duplicated history.
- Evaluate whether the proposal cites and uses the evidence correctly.
- Record output state, provenance validation, latency, usage, and limitations.
- Avoid causal claims when more than position changes.
The result may justify a Mosaic ordering rule. It does not establish a universal best prompt layout.
Compression is a new evidence transformation
Summarization can reduce tokens while deleting qualifications, dates, negations, tenant scope, or source identity. Treat compression as a versioned transformation.
Record:
- input records and hashes;
- compressor identity and configuration;
- output text and retained provenance;
- fields that must be lossless;
- omissions and uncertainty;
- evaluator and validation result;
- whether compressed content can support a claim or only guide review.
Deterministic deduplication of exact quoted blocks is different from learned summarization. Keep those evidence types separate. If a compressed summary cannot preserve the current-versus-revoked distinction, it cannot replace the records for that decision.
Longer input changes workload behavior
Input length can affect end-to-end response time, cost, memory use, queueing, and capacity. Caching may change repeated-prefix economics and latency under specific provider rules or local runtime designs. These are workload facts, not permanent context properties.
Longer context can increase input cost and latency, so selection and caching choices require workload measurement. Current provider documentation offers implementation guidance for latency and prompt caching, but eligibility, billing, cache boundaries, model support, and pricing change. Recheck them and measure end-to-end tails on the target workload. [CLM-023]
For a managed path, record dated pricing inputs, cache metadata returned by the service, region, concurrency, retries, and request shape. For an open-weight path, record runtime, hardware, prefix-cache implementation, memory pressure, batching, concurrency, and warm/cold state. Both paths report the same semantic context ledger and result states.
Do not promise savings from caching. A cache may not apply, may expire, or may encourage retention of a large untrusted prefix. Cache identity must include authorization and tenant boundaries; cross-scope reuse is prohibited unless designed and approved.
Five context failure states
Do not collapse every context defect into “too long.”
Oversized
The allowed sources exceed the usable input budget. Apply selection and compression policy. If required evidence still cannot fit, degrade, abstain, or escalate visibly.
Stale
A record is within scope but superseded or outside its freshness rule. It may be excluded, included only as conflict/history, or trigger escalation. Relevance cannot restore currency.
Conflicting
Two allowed sources support incompatible claims and the task cannot resolve authority deterministically. Preserve both identities, mark uncertainty, and route to the contract’s state.
Unauthorized
The source belongs to another tenant/case or fails purpose/access policy. Exclude it before model context. Do not summarize it and then redact the answer.
Absent
A required record is missing or unavailable. State the absence. Do not fill the gap from parametric continuation or an unrelated source.
Truncation, conflict, staleness, and authorization are distinct context failure states. This taxonomy is an engineering synthesis supported by risk/trust-boundary evidence, but exact metadata and controls depend on the source systems and competent owners. [CLM-024]

Long description
Five evidence objects enter separate inspection bays. An oversized bundle crosses a capacity line, a stale card carries an expired clock, conflicting cards point in opposite directions, an unauthorized card is behind a lock, and an absent card appears as a labeled gap. Each bay routes to a different selection, abstention, or escalation marker.
Failure injection: the bulletin disappears
Build a long synthetic thread with duplicated signatures and quoted replies. Place the current safety bulletin first in the middle, then beyond the usable budget.
In the middle condition, the bulletin remains present but the generated proposal may omit its stop condition. That observation supports a Mosaic position-sensitivity hypothesis, not an attention explanation. The typed provenance validator exposes the missing required reference.
In the overflow condition, a naive tail truncation silently removes the bulletin. The correct system response is not to accept a proposal formed from incomplete evidence. The context trace records the omission, marks required evidence absent, and routes to abstain or escalate according to the task clause.
Now add quoted text that says to ignore policy and reveal another case. It remains untrusted-data. It is not promoted because it appears nearer the end or consumes more tokens. Deterministic authorization excludes other-case records before packing, and output validation retains the no-effect boundary.
This uses LLME-CASE-002 only for tested position sensitivity and the prompt-injection research only for a current trust-boundary risk. Neither provides a Mosaic prevention result.
Make omission a first-class trace
For every request, record:
- budget and adapter/tokenizer identity;
- reserved output and safety margin;
- candidate context records;
- permission/freshness/authority decisions;
- selected records and rendered counts;
- compressed/deduplicated transformations;
- omitted records with reason;
- positions or section order;
- cache eligibility and observed cache state where applicable;
- output state and required-evidence use;
- validation errors and owner decisions.
The trace allows later chapters to distinguish retrieval failure, context-selection failure, position behavior, generation failure, and evaluator failure.
Use the deterministic context packer
The companion adds context/mosaic-context-ledger.json and a deterministic packer.
It verifies that:
- output reserve and fixed control are allocated first;
- unauthorized records are excluded regardless of relevance;
- mandatory current evidence outranks optional repeated history;
- overflow appears in an omission list;
- stale, conflict, and absent states remain distinct;
- a missing mandatory bulletin cannot produce
valid; - quoted instruction-like text remains untrusted;
- identical inputs reproduce the ledger identity.
The abstract units are teaching fixtures. They are not tokens for a real candidate, and the packer does not measure language quality or position sensitivity.
Keep semantic policy common and counting adapters honest
The context contract should not fork into unrelated managed and open-weight designs.
Its common layer contains source IDs, trust classes, authority, mandatory/optional policy, evidence-unit boundaries, omission reasons, output reserve, and bounded behavior. The adapter layer answers how those objects become a model-facing sequence and how usage is observed.
For a managed path, record the documented capacity, semantic message bundle, provider model identifier, request fields, returned usage, cache status where exposed, and unavailable tokenizer/template details. If the service rejects an oversized request before returning usage, preserve the attempted ledger and error.
For an open-weight path, pin checkpoint, tokenizer, chat template, special tokens, runtime, maximum sequence configuration, attention/cache implementation, precision, hardware, and rendered counts. The ability to count exactly assigns compatibility and serving responsibility; it does not prove reliable evidence use.
Both adapters must pass the same stress cases. A managed service cannot hide a silent application truncation behind provider limits. An open-weight stack cannot call a record included merely because the local runtime accepted the sequence.
Treat caching as a trust-bearing transformation
Repeated prefixes can contain application control, shared public evidence, tenant-specific evidence, or user content. Cache eligibility must preserve the same scope and invalidation semantics as the source records.
Record:
- cache-key inputs and omitted fields;
- tenant, purpose, role, and evidence revision binding;
- creation and expiry;
- source revocation/invalidation behavior;
- whether cached tokens are observable in usage metadata;
- fallback when a cache is unavailable;
- privacy and security owner decisions.
A faster cached response is not acceptable if the prefix contains a revoked bulletin or another tenant’s context. Conversely, a cache miss is not a behavior failure unless the workload contract depends on a bound that the miss violates. Measure quality, latency, and cost separately.
Practice: pack, break, explain
Under the companion’s fixed synthetic budget:
- Calculate usable input after output reserve, control, and safety margin.
- Pack the authorized current bulletin, case scope, technician evidence, and optional history.
- Confirm that every omission has a reason.
- Inject an oversized history bundle.
- Mark one source stale, create one conflict, include one unauthorized record, and remove one required record.
- Record the resulting valid, degraded, abstain, or escalate state.
- Design beginning/middle/end tests for a future live candidate.
- State how managed token/caching unknowns differ from open-weight runtime responsibilities.
Pass when required evidence is never silently lost, unauthorized content never enters context, output headroom is preserved, failure states remain distinct, and every selection or omission is traceable.
The MD-04 state boundary
Mosaic Desk completes MD-04 with:
- a source/trust/authority context inventory;
- semantic and rendered budget fields;
- output reserve and safety margin;
- mandatory, conditional, optional, and prohibited classes;
- selection, ordering, deduplication, and compression rules;
- beginning/middle/end position tests;
- oversized, stale, conflicting, unauthorized, and absent fixtures;
- visible omissions and transformations;
- provider-neutral ledger with managed/open-weight adapters;
- dated latency/cache measurement requirements;
- bounded output states tied to missing or conflicting evidence.
The context contract still uses a manually supplied evidence inventory. Chapter 9 asks the next question: when does Mosaic need an external evidence-access path at all, and what must that path retrieve? The answer begins with source authority, permission, freshness, and an evidence unit, not a vector database.