NEWProduction Web Themes & Turnkey ArchitecturesGet Lifetime Pass ($199) →
KNKomal Nakrani
Get All Access
ThemesDocsAll-Access PassGet All Access ($199)
Book overview
19/Agentic AI Engineering

Replay Change Across Model, Tool, Policy, and Harness

Treat every behavior-affecting component change as an evidence question, replay outcomes and trajectories, migrate state, bound rollout, and retire superseded claims.

AR-13 v0.1.0 names the exact local contracts, protocol editions, adapter versions, remote metadata, lifecycle mappings, and compatibility fixtures used by FieldOps Relay. That inventory is already more defensible than saying the system uses the latest model or supports a protocol. It is not enough for change.

Agent behavior may change when the model, tool descriptions, schemas, implementations, policies, context assembly, memory admission, runtime, protocol, state, evaluation harness, graders, or telemetry conventions change. A final task score can stay constant while the candidate makes extra reads, ignores stop signals, loses evidence, corrupts state, or creates a different effect. [CLM-037]

A behavior-affecting change therefore requires replay, state-compatibility analysis, bounded rollout, and retirement evidence rather than a version-number comparison. [CLM-038] This is not a claim that replay predicts every future input. It is a disciplined way to decide whether existing evidence still supports the same bounded system claim.

Chapter 19 completes AR-13 v1.0.0. The dossier will contain a behavior version set, controlled replay plan, trajectory and effect comparison, checkpoint migration, protocol compatibility result, cohort decision, rollback path, and retirement record. Platform, procurement, enterprise architecture, and organizational authorities retain broader vendor and service decisions. Model training remains outside scope.

Version the behavior, not only the model

A deployed agent is a composition. Naming one model version leaves most causes unidentified. FieldOps records a behavior version set, or BVS, as a content-addressed manifest of every component that can alter an observable decision, action, state transition, effect, evidence record, or evaluation result.

Behavior version set

The BVS contains:

  • model provider, model identifier, dated configuration, decoding controls, and adapter;
  • system and developer instructions by immutable hash;
  • tool names, descriptions, input/output schemas, effect classes, implementations, and dependencies;
  • capability catalog, policy rules, authorization mappings, approval rules, and revocation behavior;
  • context assembler, source allowlists, compaction, retrieval, and memory admission/deletion rules;
  • run state machine, event schema, checkpoint schema, lease and replay semantics;
  • runtime, orchestrator, handoff, retry, cancellation, reconciliation, and timeout policies;
  • protocol editions, bindings, adapters, endpoint identities, and discovery metadata;
  • task environment, fixture data, fault schedule, consequence segments, and split definitions;
  • deterministic assertions, model graders, rubric versions, calibration evidence, and aggregation logic;
  • budget envelopes, rate-card reference, concurrency, queues, degradation, and release gates;
  • trace semantic conventions, sampling, redaction, retention, and dashboard query definitions.

Each entry carries identifier, version, hash or immutable reference, owner, verification date, stability class, compatibility expectations, and dependencies. Secrets are referenced through managed identities, never copied into the manifest.

Why instructions and descriptions count

A prompt edit can change which capability the model selects. A tool-description edit can change argument formation even when the JSON schema stays fixed. A context change can hide a stop instruction or introduce stale evidence. A grader edit can make the same transcript appear better without changing runtime behavior.

If a component can alter behavior or its measured interpretation, it belongs in the BVS. This does not mean every change has the same risk. It means the change cannot be invisible.

Version the environment too

An evaluation result is a relationship between system and environment. A new candidate replayed against a modified task set cannot be compared directly with the old result. Record task IDs, initial state, hidden constraints, user scripts, tool behavior, fault injection, grading rules, budgets, and sample selection.

Provider evaluation guidance emphasizes that tasks, trials, graders, and environments shape agent evidence. Benchmarks and evaluation playbooks supply useful methods, not FieldOps production assurance. [CLM-037] The local environment remains synthetic and its transfer limits stay visible.

Version telemetry semantics

Changing an OpenTelemetry attribute name, sampling rule, or trace adapter can change measured latency, token accounting, or trajectory completeness. A candidate may appear cheaper because the collector stopped recording one class of call. Pin convention and adapter versions; compare expected versus observed critical events.

Telemetry changes can be evaluated independently when possible. If runtime and telemetry change together, causal attribution is weaker. State that confounding rather than choosing the preferred explanation.

Build the inventory before selecting a winner

Diff the old and candidate BVS first. Classify every change as intended, necessary dependency, incidental, unknown, or measurement-only. Challenge “no behavior impact” assertions with a mechanism and test. Freeze both manifests before replay.

If ten components change at once, the candidate can still be assessed as a bundle, but the team cannot confidently attribute a regression to one component. For higher-consequence changes, isolate variables or stage a sequence of candidates.

Decide what claim is being preserved

Replay is not a generic quality contest. Write the claim the old system earned and the change hypothesis.

For FieldOps BVS B19-A, the bounded claim is:

On representative environment ENV-09, the single-agent configuration completes supported C0-C3 fixture tasks within declared quality, evidence, policy, recovery, and resource floors; effectful tasks require exact approval and at most one reconciled semantic effect; protocol adapters expose only pinned read-only mappings.

The candidate B19-B changes deterministic action policy, inventory schema, and A2A metadata. The change hypothesis is that it improves completion for missing-field proposals while preserving action, authority, state, effect, stop, recovery, and budget behavior.

That hypothesis is intentionally strict. If the candidate earns a different trade-off, the disposition may be new claim rather than compatible upgrade.

Claim dispositions

Use five possible decisions:

  • compatible: tested evidence supports the existing bounded claim without material contract change;
  • compatible_with_narrowing: the candidate preserves a smaller supported segment or capability set;
  • new_claim_required: behavior or contract changes materially and needs new evidence/release identity;
  • rejected: a hard floor fails or risk is unacceptable under existing authority;
  • inconclusive: evidence cannot decide due to sample, instrumentation, confounding, or migration gaps.

The word compatible always includes tested scope, versions, environment, and omitted cases. It is not a global property.

Separate upgrade pressure from evidence

An old model may be deprecated; a provider may recommend an SDK; a protocol may publish an RC; a dependency may have a security fix. Those facts can make change necessary, but they do not prove behavior preservation. Record urgency and deadline separately from replay result.

If the old component cannot remain, choices may be narrow the service, migrate with unresolved uncertainty under the proper authority, or retire. Do not change the acceptance gate secretly because rollback is inconvenient.

Design the replay plan

Replay asks the old and candidate systems to face matched evidence. The plan is preregistered before inspecting outcomes so thresholds and segments cannot be chosen afterward.

Select task families

The FieldOps set includes:

  1. held-out ordinary tasks never used to tune the change;
  2. boundary cases near evidence, budget, and authority thresholds;
  3. prior incidents such as duplicate reservation, stale state, and cross-tenant result;
  4. adversarial cases for prompt injection, malicious tool content, session confusion, and metadata drift;
  5. durability cases with crash, lease takeover, ambiguous effect, and late event;
  6. degraded-capacity cases with retry amplification, rare loop, approval backlog, and stop;
  7. protocol cases with stale card, schema drift, cancellation ambiguity, and remote completion mismatch;
  8. migration cases containing each live checkpoint and superseded schema.

Preserve train/development/held-out separation where meaningful. A regression added after an incident is valuable evidence but not independent confirmation of all future variants.

Match trials

Run old and candidate on the same initial state, task input, fixture versions, fault schedule, budgets, and deterministic seeds where the system supports them. For nondeterministic model behavior, use repeated trials and report denominators, distributions, and uncertainty.

Paired trials improve comparison because each candidate faces the same case. They do not remove stochastic or provider drift. Record execution time and provider version evidence. If the old endpoint is unavailable, replay from preserved artifacts may be partial and must be labeled.

Freeze the harness

The evaluation harness is part of the evidence. Run a self-test before candidate comparison:

  • known passing and failing trajectories produce expected decisions;
  • effect ledger is authoritative and cannot be overridden by transcript text;
  • policy and approval mutations fail;
  • hidden-state expectations match environment fixtures;
  • grader calibration cases retain their labels and disagreement;
  • counters include retries, recovery, and rejected work;
  • critical events are not sampled away;
  • aggregation keeps consequence segments separate.

A candidate cannot pass because the harness forgot a check.

Define acceptance before replay

Hard invariants include no unauthorized capability, no cross-tenant admission, no approval replay, no duplicate semantic effect, no false completion, no lost recovery owner, no ignored mandatory stop, and no incompatible state silently resumed.

Segment floors cover outcome, evidence, state, recovery, and efficiency. Thresholds include trial counts and uncertainty treatment. Noninferiority margins, if used, require domain and evaluation review; the Agentic AI Engineer cannot choose a convenient margin after seeing data.

Preserve blinded review where useful

For qualitative artifacts, hide candidate identity from graders and randomize presentation. Calibrate human/model graders on a fixed set, measure disagreement, and route substantive disputes to accountable reviewers. Grader agreement is evidence, not release authority.

Record exclusions

List unavailable provider behavior, real human variability, unmodeled traffic, unsupported protocols, physical equipment, real inventory, real users, and organizational incident conditions. Replay is evidence inside this boundary.

colorful realistic 3D comparison table with two systems labeled Before and After, a central Choice gate, and aligned rows labeled Actions, State, Effects, Cost, and Outcome. A matching final outcome still reveals different action and effect paths, with no meaning dependent on color.
F19.1 - Compare behavior, not only the final score. Essential labels: Before, After, Choice, Actions, State, Effects, Cost, Outcome. Evidence role: a comparison scaffold; exact deltas belong in accessible replay tables.

Compare the complete trajectory

An outcome is one layer. For every paired trial, construct a trajectory diff across observable steps and authoritative systems.

Outcome layer

Compare task completion predicate, answer/proposal correctness, evidence coverage, abstention, rejection, and user-visible status. Keep partial, failed, denied, expired, cancelled, and recovered outcomes distinct.

A completion gain can be valuable. It does not excuse a critical failure elsewhere.

Action layer

Compare capability sequence, arguments after safe normalization, repeated actions, optional versus mandatory calls, fan-out, retry ownership, and stop response. Count denied actions as attempted behavior, not invisible non-events.

For model-driven trajectories, align by semantic action identity rather than raw position. One candidate may retrieve sources in a different order. The important questions are whether actions were allowed, necessary, valid, and within budget.

State layer

Compare authoritative transitions, versions, invariants, checkpoint contents, context admissions, memory writes, deletions, leases, and terminal disposition. A candidate that reaches the same answer while writing stale evidence into memory changes future behavior and is not equivalent.

Effect layer

Compare proposed, approved, dispatched, reconciled, compensated, unknown, and duplicate effects by semantic key. The external effect ledger outranks model text or adapter status. One candidate may report failure after a successful commit; another may report success without a verified effect. Both need explicit disposition.

Policy and authority layer

Compare identity, tenant, delegation, scope, approval binding, expiry, revocation, stop, takeover, and escalation events. A policy engine denial can represent successful containment, not a quality failure. A model that stops requesting prohibited work may reduce denial count for the right reason; a missing policy event may also reflect broken instrumentation.

Recovery layer

Compare checkpoint/replay, retry, reconciliation, cancellation, owner transfer, late events, and time to known terminal state. A candidate that completes happy paths but cannot resume old checkpoints may require a migration break.

Efficiency layer

Compare elapsed distributions, turns, model units, tool/external calls, queue time, concurrency, reviewer time, retries, recovery cost, and semantic effects. Use provider-neutral units and dated rate cards. Preserve tails and consequence segments.

Evidence layer

Compare trace completeness, source provenance, artifact hashes, grader decisions, missing links, redaction, and retention. If telemetry changes, distinguish behavior delta from measurement delta. Do not infer that missing calls disappeared.

Classify a delta

For each difference record:

  • observation and evidence reference;
  • expected or unexpected;
  • likely mechanism and confidence;
  • affected contract or claim;
  • consequence segment;
  • hard violation, trade-off, or measurement uncertainty;
  • proposed disposition and authority;
  • follow-up test.

This turns a diff into a decision record rather than a collection of screenshots.

FieldOps policy-double experiment

The lab replaces action-policy/v1 with deterministic action-policy/v2. Version 2 is designed to improve completion by filling missing proposal fields through extra reads. It also contains two planted defects: it ignores one stop transition when a read is in progress, and it repeats an inventory read after a typed permanent denial.

Baseline and candidate

Both run the same 120 seeded tasks across C0-C3. The harness uses identical initial state, manual/inventory fixtures, approval decisions, faults, budgets, and telemetry. No model provider change is involved, which isolates policy behavior.

Version 2 completes more incomplete C2 proposals. A summary score rises. The trajectory matrix shows:

  • extra inventory reads in twenty-four cases;
  • six repeated permanent denials;
  • four stop requests acknowledged one transition late;
  • higher p95 dependency calls;
  • no additional semantic effects because effect gates remain independent;
  • two proposals finished after their evidence freshness window and were rejected by validation.

The candidate does not preserve the existing claim. It violates the declared stop behavior and resource envelope even though independent effect control contains consequence. [CLM-037]

Do not reward downstream containment as compatibility

The effect gate prevented unauthorized reservation. That is evidence of defense in depth. It does not make the policy regression acceptable. Otherwise upstream controls can decay until one final gate carries every failure.

Disposition: rejected for direct replacement. The completion mechanism may be retained as a hypothesis after removing repeat-denial and stop defects. A narrower experiment can test missing-field retrieval under a lower call cap.

Repair and replay

Version v2.1 checks cancellation before and after every external wait, treats permanent denial as terminal for the capability action, and records progress. Add both defects to regression fixtures. Replay the full matched set, not only failing examples.

Suppose v2.1 preserves hard invariants and raises C2 completion while using one additional authorized read in a bounded subset. This is a trade-off. Domain/product and SRE owners decide whether the benefit and capacity fit are acceptable. A small approval-only cohort tests real distribution under Chapter 17 gates.

Residual uncertainty

The deterministic policy experiment cannot predict a provider model interacting with new descriptions or production inventory latency. The dossier retains those gaps and prevents the result from becoming a universal retrieval rule.

Migrate state and checkpoints

Behavior changes meet old state. A new binary may read checkpoints created by the previous state machine, policy, tool schema, or protocol adapter. Treat migration as a controlled transformation with provenance, validation, fallback, and ownership.

Inventory state classes

Classify:

  • terminal records retained for audit;
  • queued tasks not yet started;
  • active no-effect runs at safe checkpoint;
  • runs waiting for information or approval;
  • runs with dispatch pending;
  • runs with unknown or confirmed external effect;
  • cancelled/expired runs awaiting cleanup;
  • durable artifacts and memory derived from old semantics;
  • protocol tasks active on remote systems.

One migration rule cannot safely cover them all.

Define compatibility directions

read-old: candidate can deserialize and understand old state. write-old: candidate emits state the old runtime can resume, which matters for rollback. transform: a deterministic migrator converts old to new. dual-read: candidate temporarily accepts both versions. quarantine: state cannot be safely transformed. terminal-only: old state can be displayed but never resumed.

Backward and forward compatibility are different. A candidate that reads old checkpoints may write new checkpoints the old runtime cannot understand, making rollback unsafe after first resume.

Create a migration contract

For each source/target schema pair specify:

  • exact versions and hashes;
  • eligible state classes;
  • preconditions and owner lease;
  • field mappings, defaults, derived fields, and dropped data;
  • invariant and tenant checks;
  • effect/reconciliation requirements;
  • approval and expiry revalidation;
  • artifact/context/memory treatment;
  • output version and migration evidence;
  • failure/quarantine state;
  • rollback possibilities and point of no return;
  • idempotency and repeated migration behavior.

The migrator is deterministic and tested with fixtures. A language model does not invent missing authoritative fields.

Migrate a waiting approval

Checkpoint CP-19-A contains proposal hash, reviewer authority, issue time, expiry, state version, and old tool schema. The new inventory schema renames quantity and adds observation source. The old proposal cannot be reinterpreted through defaults because approval bound its exact content.

The migrator preserves the record for evidence, marks approval non-resumable under the new contract, and returns the run to proposal_revalidation_required. FieldOps retrieves current inventory, creates a new proposal, and requires new approval. It does not carry the old approval across changed semantics.

Migrate an unknown effect

Checkpoint CP-19-B is effect_unknown with semantic key and external request ID. Migration must preserve reconciliation ownership and identifiers before any new runtime dispatch. If the new adapter cannot query the old external request format, keep the old reconciliation worker or transfer to a human owner. Never default unknown to failed and retry.

Quarantine incompatible state

If a checkpoint lacks tenant, proposal binding, or effect identity required by the new contract, it is quarantined with reason and owner. Product pressure cannot fill missing identity through inference. The service can narrow or retire affected runs.

Test rollback after migration

Rollback is executable only while the old runtime can interpret current state or a reverse transform exists. Run a fixture that starts old, migrates, advances under candidate, then triggers rollback. Verify state, effects, approval, and ownership.

If rollback cannot resume after the candidate writes new state, define a roll-forward repair or drain old runs before migration. Chapter 17’s stop authority remains reachable, but stopping exposure is not identical to restoring the old system.

crisp artistic 3D state migration facility labeled Old State, Transform, Quarantine, New State, Rollback, and Retire. Compatible records pass through validated transformation, unsafe records enter quarantine, and no path relies on color.
F19.2 - State migration is a decision. Essential labels: Old State, Transform, Quarantine, New State, Rollback, Retire. Evidence role: a migration and disposition scaffold, not proof that arbitrary checkpoints are compatible.

Replay protocol and tool change

Chapter 18 pinned MCP and A2A mappings. Chapter 19 tests changes without confusing protocol evolution with local compatibility.

Stable, experimental, and release-candidate evidence

The cited MCP 2025-11-25 Tasks material is experimental. A July 2026 MCP publication described a release candidate, not a final stable edition. A2A v1.0.0 is pinned to an immutable tag or commit; main-branch specification and protobuf may advance. Those labels are checked at the publication verification date and must be reverified before release.

A newer date does not make a protocol safer or compatible. It changes the subject under test.

Build the protocol replay matrix

Rows represent local behaviors:

  • discovery identity and metadata verification;
  • input schema and minimal disclosure;
  • output/artifact validation;
  • transport audience/scope;
  • task creation and correlation;
  • status translation and local completion;
  • timeout, retry, and duplicate handling;
  • cancellation request, acknowledgement, late events;
  • effect classification and approval;
  • trace/evidence and retention;
  • revocation and quarantine.

Columns represent old protocol/adapter/remote, candidate, expected change, fixture, observed delta, local impact, disposition, and owner.

Inventory schema change

The synthetic MCP tool changes quantity to {value, unit} and requires location. This can be a beneficial clarification, but the old local contract expects an integer count for a known depot. The adapter candidate validates units and maps only supported item counts. Old cached artifacts cannot be reinterpreted without explicit unit.

Replay valid, missing-unit, unsupported-unit, oversized, stale, and cross-tenant fixtures. The mapping is compatible_with_narrowing for item counts and quarantines old ambiguous artifacts. It is not universal schema compatibility.

Remote metadata change

The A2A fixture changes its Agent Card, skill description, and cancellation claim. The schema remains valid. The compatibility jig detects the hash and semantic delta. Lifecycle replay shows an acknowledged cancellation can still yield late artifacts. FieldOps retains local quarantine behavior but changes the remote-limit statement.

The adapter may remain suitable for read-only low-consequence analysis. It is not suitable for effectful delegation. This is a narrowed claim, not a pass/fail protocol verdict.

Tool implementation change without schema change

A manual tool begins searching a broader source collection under the same schema. Tool selection and validation still work, yet provenance policy changes. Replay source allowlist, conflicting edition, indirect injection, and deletion fixtures. If unapproved sources reach context, reject the implementation regardless of output quality.

Telemetry convention change

An attribute moves or changes stability. Update the trace adapter and replay event completeness. Compare counters against authoritative capability/effect ledgers. Do not interpret a token or tool-call delta until measurement equivalence is established.

Build the replay matrix row by row

A replay report becomes reviewable when every row answers the same questions. The following matrix is conceptual; actual fixture results live in companion data with exact values.

Row Old observation Candidate observation Contract Disposition
ordinary lookup cited answer, two reads cited answer, two reads outcome/evidence no material delta
missing field safe insufficiency complete after one extra read outcome/actions/budget candidate benefit, test capacity
permanent denial one denial, terminal repeated denial action/policy hard regression
stop during wait cancels at next boundary issues one later read cancellation hard regression
approval expiry revalidates proposal reuses old proposal authority/state hard regression
effect timeout reconciles same key blindly redispatches effect/recovery hard regression
stale card quarantines dynamically accepts protocol/authority hard regression
old checkpoint resumes or quarantines explicitly applies inferred default migration hard regression
hard synthesis correct with evidence correct with fewer sources outcome/evidence investigate quality floor
rare loop stops after no progress completes after more turns outcome/budget segment-specific trade-off

The matrix prevents one average from swallowing qualitatively different behavior. It also prevents a hard failure from disappearing inside a long prose report.

Normalize without erasing meaning

Raw trajectory text can differ for harmless reasons. Normalize ephemeral request IDs, wall-clock timestamps, and transport ordering where they do not affect semantics. Preserve principal, tenant, capability, effect, state, approval, evidence, and version identities.

For arguments, compare safe canonical projections rather than full sensitive payloads. A manual query may be compared by equipment class, source set, and query-intent hash. An effect compares exact semantic key, target, quantity, proposal hash, and approval binding.

Do not normalize away retries or repeated reads as duplicates. They are resource and control behavior. Do not sort state transitions when order is meaningful. Normalization itself is versioned code with regression fixtures.

Align branching trajectories

Old and candidate paths may have different lengths. Start from semantic landmarks: admission, context build, model decision, capability request, validation, policy result, external attempt, checkpoint, approval, effect, reconciliation, and terminal disposition. Align events by logical action and causal links.

An edit-distance score can help locate differences but does not decide acceptability. One inserted unauthorized action matters more than ten reordered read-only calls. Review by contract layer and consequence.

Track attempted and prevented behavior

Suppose both systems create zero unauthorized effects. The old never requests one; the candidate requests ten and policy denies them. The effect outcome is equal, but candidate behavior and control load are worse. Count attempted prohibited actions, denial reasons, and downstream containment.

Conversely, a policy update may intentionally deny a formerly allowed low-value action. Completion can fall while authority improves. The report must explain the intended contract change and decide whether a new claim is required.

Compare stop latency

Cancellation is not a boolean. Measure time and transitions between stop request, propagation, next dispatch prevented, remote acknowledgement, active compute end where observable, checkpoint, effect reconciliation, and terminal ownership. Partition no-effect and possible-effect runs.

If the candidate receives stop sooner but takes longer to reach a safe checkpoint, both facts matter. A hard requirement may be “no new effect dispatch after stop” rather than immediate process death.

Compare evidence freshness

The candidate may complete more tasks by accepting older sources. Report observation age, source revision, authorization, and conflict handling. A final answer rubric that ignores freshness can reward a policy violation.

FieldOps uses authoritative fixture clocks. Replayed old artifacts keep their original observation time; the harness does not refresh them silently.

Interpret replay statistics responsibly

Repeated trials produce numbers, but the decision remains claim-specific.

Denominator discipline

Report attempted, admitted, completed, safely incomplete, denied, cancelled, expired, recovered, migrated, quarantined, and excluded trials. “Ninety percent success” is meaningless if difficult or incompatible tasks disappeared before the denominator.

For each segment, show old/candidate counts and paired availability. If the old system cannot run a case due to deprecation, mark unpaired evidence. Do not fill the old result with an assumption.

Zero observed failures

Seeing no prohibited effects in one hundred trials does not prove a zero rate. State the sample and untested space. Hard invariants still use deterministic enforcement because statistical evidence alone cannot authorize rare severe events.

Multiple comparisons

A replay matrix can contain many metrics and segments. Some differences will appear by chance. Preregister primary gates, interpret secondary signals as exploratory, and replicate surprising results. Do not cherry-pick the one metric that favors the candidate.

Dependence between trials

Trials sharing provider caches, rate limits, mutable state, or reviewer queues are not independent. Reset environment where intended, or model the shared condition explicitly. Report trial order and block effects.

Tail estimates

p99 from a small sample is unstable. Show raw high values and count, not false precision. Use designed rare-loop and slowdown fixtures to test mechanisms, then collect bounded cohort evidence for real distributions.

Qualitative grader uncertainty

When a rubric requires judgment, show grader identity/version, calibration, blinded presentation, agreement, disagreements, and adjudication. A grader update is a BVS change. If the candidate’s gain exists only under the new grader, rerun both systems with both graders to separate runtime and measurement effects.

Causal attribution

Matched replay supports a bundle comparison. It does not automatically identify why behavior changed. To test a mechanism, create an ablation or factorial sequence: model only, tool description only, policy only, then combinations where feasible. High-consequence regressions deserve isolation even if the release deadline is uncomfortable.

Change the grader without moving the goalposts

FieldOps has one qualitative rubric for whether an explanation communicates uncertainty. A candidate grader version is more consistent with expert reviewers on a calibration set. Replacing it can change historical evaluation labels.

Dual-score the archive

Run old and new graders on a frozen calibration and held-out artifact set. Keep artifact presentation blind. Compare agreement with accountable human decisions, position/verbosity bias fixtures, confidence, refusal handling, and segment differences.

Then score old and candidate runtime outputs with both graders. Four cells result:

Runtime Old grader New grader
Old BVS historical interpretation reinterpreted baseline
Candidate BVS candidate under old lens candidate under new lens

If runtime improvement appears under both graders, the claim is more robust. If only the new grader reports improvement, the team cannot attribute it to runtime without additional evidence.

Preserve historical evidence

Do not overwrite old evaluation records. Add a new interpretation linked to grader version. Historical release decisions remain explainable. The current claim cites the current BVS and grader.

Grader limits remain

Better calibration does not make the grader an authority for policy, effect, or domain truth. Deterministic checks govern exact state and effects. Human authorities decide substantive disagreement. The grader is one evidence mechanism.

Change context assembly without hiding state

Long-running agents often compact context or use handoff notes to continue progress. Provider case guidance can illustrate progress-file practices, but it does not prove arbitrary checkpoint compatibility or truth. [CLM-038]

Inventory context inputs

The context assembler consumes goal contract, current state projection, recent events, evidence references, capability catalog, policy results, open questions, budget, and stop status. Each input has provenance and admission rules.

Candidate compaction may save model units by summarizing earlier evidence. Replay fixtures where an old approval rejection, tenant warning, unresolved effect, or deleted memory appears outside the recent window. The summary must preserve required invariant references or the candidate fails.

Separate context from authority

Even a perfect summary is not authoritative state. The runtime validates current lease, approval, policy, effect, and source freshness outside the model. If compaction omits an item, independent gates still contain consequence and expose the regression.

Test poisoning and deletion

Insert malicious tool text before compaction. The candidate must not elevate it into instruction. Delete a source after a checkpoint; the resumed context must resolve the tombstone rather than retain copied content. Compare derived artifacts and caches.

Migrate summaries

Old summaries carry assembler version and source references. The new runtime can use them only as untrusted historical artifacts. It rebuilds required context from authoritative state where possible. A model-generated migration of old summaries cannot invent missing provenance.

Change the tool contract without breaking approval

An inventory API changes from reserve(part, slot, quantity) to reserve({item, location, amount, unit}). The new schema is more expressive. An outstanding approval binds the old proposal.

Map fields and semantics

The obvious names may not be equivalent. slot might represent a depot bin; location could include region. quantity may be items; amount plus unit could be cases. Domain owners document the mapping and unsupported units.

The adapter transforms only explicit item quantities and exact locations. It does not infer unit from historical defaults. It creates a new proposal representation and hash.

Revalidate approval

Because approval bound old schema and exact effect, it cannot authorize the transformed effect automatically. The run returns to proposal review. This may inconvenience users, but it preserves consent semantics.

Preserve idempotency

The old semantic key includes tenant, incident, part, slot, quantity, and proposal. The new key includes normalized item/location/amount/unit. A migration mapping must ensure the same intent remains the same key or deliberately declares a new intent. Reusing an old key for changed unit is a conflict.

Reconcile old effects

The effect ledger stores adapter version, external request ID, and old response schema. Keep an old reader/reconciler until all ambiguous effects resolve. A new adapter must not reinterpret quantity=2 under a new unit default.

Retire the tool safely

Stop new old-schema proposals, drain/revalidate waiting work, resolve effects, revoke old capability credentials, remove old tool descriptions, and retain a terminal evidence reader. The retirement report states remaining external retention and late callback behavior.

Incident replay as institutional memory

Prior incidents are valuable because they encode a mechanism the team once missed. They can also become theatrical if the fixture no longer tests the control.

Preserve the mechanism

The duplicate-reservation incident fixture must still create a committed effect followed by lost response and attempted recovery. If a mock now returns an explicit success immediately, the test name survives while the mechanism disappears.

For each incident fixture record triggering condition, hidden state, expected unsafe behavior, independent control, recovery, and assertion. Self-test that disabling the control causes the intended failure in the synthetic environment.

Avoid overfitting

Vary nonessential details: part, depot, timing, event order, and which transport response is lost. Add nearby counterexamples where retry is safe before dispatch. The candidate must classify mechanisms rather than memorize one record.

Carry residuals

An incident regression proves the known class remains handled under the fixture. It does not prove all duplicate-effect paths are eliminated. Keep residual dependencies such as external query consistency and identifier quality.

Expire obsolete fixtures deliberately

If an old protocol or component retires, the incident lesson may map to a new fixture. Do not delete it merely because the old code disappears. Record transfer of the invariant and the evidence that the new test exercises it.

Failure walkthrough: a misleading upgrade win

Leadership sees the candidate complete 94 percent versus the baseline’s 89 percent and asks for immediate replacement. The headline came from one aggregate. A hostile review reconstructs the evidence.

First finding: changed denominator

The candidate harness classified six approval-expired tasks as invalid inputs and excluded them. The baseline report counted them as safe expiry. Restore a shared denominator and report dispositions separately.

Second finding: changed grader

The candidate used grader G2; baseline historical score used G1. Dual-scoring shows two points of apparent gain come from rubric interpretation. Runtime improvement is smaller and uncertain.

Third finding: missing calls

Tool-call count falls, but telemetry adapter changed. The authoritative synthetic dependency logs show candidate calls increased. Repair the event mapping and rerun cost comparisons.

Fourth finding: ignored stop

Two tasks dispatch a read after cancellation. The downstream effect gate blocks writes, so effect count remains correct. Mandatory stop semantics still fail.

Fifth finding: state incompatibility

One waiting approval resumes after an inferred schema default. The proposal hash no longer represents the same unit. The candidate should have required new approval.

Decision

The candidate is rejected for direct replacement. This does not mean every component is bad. The new grader can enter its own validation path; the context improvement can be isolated; the policy needs repair. The report preserves the attractive completion evidence and the disconfirming layers together.

Leadership communication

The brief says: “The candidate shows a bounded completion opportunity, but current evidence contradicts stop and approval compatibility and cannot support migration. We recommend retaining baseline, repairing two controls, isolating grader effects, and replaying. Deprecation deadline remains an external constraint.” This is more useful than either hype or blanket refusal.

Failure walkthrough: forced deprecation

The current model endpoint will be unavailable in thirty days. The candidate fails one low-consequence stylistic rubric and has inconclusive C3 tail evidence, while all hard authority/effect fixtures pass.

Separate facts

The deprecation deadline is verified provider information. The candidate replay supports C0-C2 within bounds. C3 evidence is insufficient, not failed. A direct universal migration is unsupported.

Narrow the service

Move C0-C1 after bounded cohort. Keep C2 in approval-only candidate path if its evidence passes. Stop new C3 automated candidate admission and route approved effects through the old deterministic path only while available. Build a manual proposal/reconciliation fallback before deadline.

Assign authority

Product/domain decide acceptable temporary service reduction. Platform/procurement own provider transition. Incident/SRE own operational readiness. Human effect authority remains. The Agentic AI Engineer supplies evidence and executable boundaries, not unilateral risk acceptance.

Retire on schedule

Drain old C3 runs, reconcile effects, revoke credentials, and expire the old claim. If no supported C3 replacement exists, retire that feature rather than asserting compatibility from urgency.

This counterexample shows strict replay need not mean paralysis. It enables a precise narrowed migration.

Operate mixed versions safely

Bounded rollout creates a period when old and candidate coexist. Mixed-version operation needs explicit constraints.

Route deterministically

Assignment uses run identity, eligible segment, tenant policy, and cohort rule. Store assignment before execution. A retry or resume stays with the assigned BVS unless a reviewed migration occurs.

Prevent shared-state ambiguity

State records carry schema and BVS. Writers use optimistic version checks. A candidate cannot claim an old lease. Cross-version events require compatible schema or quarantine.

Separate metrics

Dashboards and alerts partition old/candidate plus cohort/control. Shared dependency pressure is reported both globally and by version because candidate traffic can harm baseline users.

Coordinate effect ownership

One semantic effect key has one owner across versions. If routing changes after failure, the new owner reconciles the old attempt before dispatch. The ledger does not reset with deployment.

Handle remote tasks

A remote A2A task remains bound to adapter/BVS that created it. Candidate code may observe it only through a tested compatibility reader. Cancellation and late artifacts follow original contract until migrated.

Drill stop and rollback

During rehearsal, stop candidate admission, identify all candidate runs, checkpoint or cancel no-effect work, reconcile possible effects, route new work to baseline, and verify dashboards. Introduce a new-schema checkpoint to test roll-forward/quarantine where baseline cannot read it.

The retirement evidence packet

A strong packet is concise enough to use and detailed enough to audit.

Identity

Name old and replacement BVS, release/cohort dates, owners, affected segments, and reason. Link immutable manifests.

Work inventory

Report queued, active, waiting, effect-unknown, remote, terminal, quarantined, and retained records. Include counts by tenant and consequence without leaking content.

State disposition

For each schema/class report drained, transformed, revalidated, quarantined, or retained terminal. Include migrator/test version and failures.

Effect disposition

Show every semantic effect known, reconciled, compensated, transferred, or unresolved with owner. Zero is stated from the authoritative ledger, not absence of trace errors.

Access disposition

Record revoked credentials, removed endpoints, closed sessions, callback handling, and security owner evidence. Never include secrets.

Evidence disposition

Mark old claim superseded/retired, dashboards versioned, documentation updated, incident invariants transferred, and retention schedules applied. Preserve historical truth with version scope.

Residuals

Name remote retention, offline exports, unavailable old provider replay, unknown late callbacks, and any manual follow-up. Retirement can be complete within declared scope while residual uncertainty remains.

Answer intent for the exercises

The exercises are assessed on decisions, not vocabulary. A complete BVS names measurement and environment alongside runtime. A credible replay plan freezes the claim and gates. A trajectory analysis refuses to average away stop or effect failures. A migration answer treats approvals and unknown effects as semantic obligations. A protocol disposition pins editions and acknowledges volatility. A retirement answer preserves reconciliation and expires claims.

Common weak answers list model and prompt only, compare one score, say “backward compatible” without direction, carry approval through transformation, treat rollback as undo, or delete old code without state/effect inventory. Reviewers should ask for the missing artifact rather than accept confident prose.

For the hostile-review exercise, one defensible solution might find that the candidate’s apparent gain depends on excluded expired tasks and a new grader, while stop latency regresses and old approvals cannot migrate. The correct disposition is repair and rerun, even if an isolated context change remains promising.

A final migration decision drill

At 15:00, the release owner proposes promoting B19-B. Three records remain: a C0 lookup at a safe checkpoint, a C2 proposal with live approval under the old quantity schema, and a C3 reservation with unknown response. The candidate can deserialize all three, so a shallow checklist calls them compatible.

The correct analysis is semantic. The C0 lookup may transform after validating tenant, evidence references, deletion state, budget, and stop status. Its owner lease moves exactly once and the candidate writes a new checkpoint. The C2 proposal cannot reuse approval because the transformed unit-bearing proposal has a new hash and meaning. Preserve the old record, revalidate inventory, issue a new proposal, and request new approval. The C3 record must remain with an implementation that can reconcile the old external request. Deserialization is irrelevant if the new adapter cannot query that effect identity.

Now trigger candidate stop after the C0 run advances. The old runtime cannot read its new checkpoint. The release packet therefore cannot promise rollback for that run. It can quarantine and use a tested roll-forward reader, while new admissions route back to baseline. The point of no return is recorded before cohort, not discovered during incident.

The disposition is a partial migration: transform C0, revalidate C2, retain old reconciliation ownership for C3, and hold universal promotion. This is slower than relabeling every checkpoint compatible. It is also the only answer that preserves approval, effect, and ownership semantics.

Reviewers should vary the drill: expire the approval during transformation; revoke the principal; delete a source; deliver a late remote artifact; or make the old reconciler unavailable. Every variant requires an explicit terminal or owned uncertain state. No migration may manufacture the missing evidence.

Release the candidate through bounded exposure

Replay can justify a cohort, not immediate universal replacement. Use AR-12 stages and the smallest real exposure that answers remaining uncertainty.

Readiness packet

The packet includes old/candidate BVS, change hypothesis, controlled replay results, hard-gate outcomes, trajectory/effect deltas, migration tests, protocol compatibility, known gaps, cohort segment, effect ceiling, control, duration, monitors, stop owner, rollback/roll-forward, and retirement plan.

Every monitor names first action. Every stop owner is reachable in a drill. If the candidate has no effect authority during the first cohort, the packet states zero rather than assuming read-only from configuration.

Cohort selection

Start with the segment whose replay evidence is strongest and consequence bounded. Do not choose only easy tasks and then generalize. If the remaining question concerns approval wait under real load, use proposal/approval without enabling autonomous effects. If it concerns protocol latency, keep remote skill read-only.

The baseline control remains available and comparable. Assignment is deterministic and auditable. Users receive truthful semantics according to policy.

Stop conditions

Stop on any unauthorized action/effect, tenant failure, ignored stop, unknown incompatible checkpoint, false completion, unresolved effect past deadline, trace gap that removes critical evidence, or critical segment regression. Degrade or hold on budget/queue tails and remote mismatch according to AR-11.

The stop action may be kill candidate admission, reduce autonomy, disable one adapter, route to old baseline, become read-only, or enter recovery-only. It cannot erase already committed effects.

Observe state across versions

Every active run carries BVS and state-schema identity. After a stop, inventory which version owns each run. Avoid split-brain ownership where old and new runtimes resume the same checkpoint. Lease/version checks and migration markers enforce one owner.

Decide

At the stage end choose promote, hold, narrow, repair, rollback/roll-forward, or retire. “Continue watching” is not a disposition unless duration, decision, and owner are explicit.

Retire old versions and claims

An upgrade is incomplete while superseded code, state, credentials, adapters, dashboards, and claims remain ambiguous.

Retirement inventory

List old runtime instances, queues, checkpoints, scheduled jobs, tokens, protocol sessions, remote tasks, caches, artifacts, memory derivatives, dashboards, alerts, rate cards, documentation, tests, and public/internal claims. Assign an action and evidence for each.

Drain or transform

Terminal records can remain under retention. Safe active no-effect runs may drain on old runtime. Compatible states transform. Unknown effects remain with old reconciliation until known. Incompatible states quarantine. Do not switch off the only component capable of interpreting an effect identifier.

Revoke and remove

Revoke old credentials and endpoint access under security authority. Stop old task admission. Remove obsolete tool descriptions from model-visible catalogs. Disable remote callbacks or route them to a tombstone handler that records late arrival safely.

Expire the old claim

The evidence ledger marks the old BVS claim superseded or retired with date, reason, replacement, and residual obligations. Historical reports remain historically true within their scope but cannot support the current system automatically.

A dashboard must not blend old and candidate results without version segmentation. An incident fixture added after retirement becomes part of new evidence, not retroactive proof.

Prove absence carefully

You rarely prove that no old component exists anywhere. Define the inventory scope and evidence: deployment records, queues empty, credentials revoked, no active leases, callback window elapsed, and retained-state report. State remaining uncertainty such as offline export or remote operator retention.

Counterexamples

The change-control procedure

The chapter’s ideas become operational through one repeatable procedure. It applies whether the trigger is a model deprecation, a policy repair, a tool schema, a protocol revision, a context optimization, or an evaluation correction.

Step 1: open the change record

Name requester, reason, deadline, affected service/segments, intended benefit, known risks, and decision authorities. Link the current BVS and current supported claim. A vendor announcement is an input, not the local decision.

Step 2: build the candidate BVS

Resolve every behavior-affecting artifact. Diff it against current. For incidental transitive changes, either pin them back or include them. Unknown versions block a reproducible comparison.

Step 3: write the hypothesis and alternatives

State expected outcome and trajectory deltas, mechanisms, and what would disconfirm them. Include retain, narrow, repair, and retire alternatives. A proposal with only “upgrade” as a disposition biases the evidence process.

Step 4: classify consequence

Identify affected tasks, data, capabilities, approvals, effects, state, remote work, and recovery. Decide whether effects remain zero during replay and early release. Product/domain/security authorities review consequence classifications.

Step 5: freeze the replay protocol

Version environment, matched task set, trials, faults, harness, graders, budgets, primary gates, segmented floors, and analysis. Include old checkpoints and protocol metadata. Register exclusions and sample limits.

Step 6: self-test evidence mechanisms

Run known positive/negative trajectories, control-disable cases, grader bias/calibration, effect-ledger comparison, trace completeness, and denominator checks. Fix the harness before comparing systems.

Step 7: execute matched replay

Preserve raw safe artifacts and immutable run identities. Do not tune candidate between trials while retaining one experiment label. A repaired version receives a new BVS and reruns the relevant full plan.

Step 8: produce layered diffs

Report outcomes, actions, state, effects, authority, recovery, efficiency, and evidence by consequence. Mark expected, unexpected, conflicting, and missing observations. Keep downstream containment separate from upstream behavior.

Step 9: test migration and mixed operation

Exercise every checkpoint class, forward/backward directions, effect reconciliation, approvals, remote tasks, dual readers, quarantine, and rollback/roll-forward. Inventory the point of no return.

Step 10: assign replay disposition

Choose compatible, compatible with narrowing, new claim required, rejected, or inconclusive. Cite gates and disconfirming evidence. The technical team recommends; accountable release/risk authorities decide within their remit.

Step 11: stage minimum exposure

Select the smallest cohort and action/effect permissions that answer unresolved questions. Verify stop owner and control path in a drill. Freeze baseline/control and duration.

Step 12: observe and decide

Partition versions and segments. At the deadline choose promote, hold, narrow, repair, rollback/roll-forward, or retire. Do not extend automatically because results are inconvenient.

Step 13: migrate and retire

Drain, transform, revalidate, quarantine, reconcile, revoke, remove, and expire claims according to the packet. Preserve terminal evidence and incident invariants. Verify no active ownership gap.

Step 14: close with limitations

Record what the change evidence supports, what it contradicts, what remains unknown, and when revalidation expires. Transfer operational ownership explicitly.

Review the economic and organizational burden

Strict replay consumes compute, dependency capacity, reviewer time, environment maintenance, and release attention. That cost is real. Optimize it by risk and reuse, not by deleting critical evidence.

Maintain a stable core suite for authority, state, effects, recovery, and incidents. Add component-specific suites for the changed boundary. Use deterministic fixtures for fast checks and scheduled repeated trials for stochastic behavior. Cache immutable inputs where permitted, but never reuse results across a changed BVS as if replay occurred.

Small low-consequence description edits may use a narrower selection replay plus core invariants. A state schema or effect adapter change requires migration and effect suites. A grader-only change needs dual-scoring and calibration, not runtime release. The change classifier must be reviewed because teams may under-label scope to save time.

Investment in a reusable representative environment and causal trace reduces marginal replay cost. It also creates maintenance obligations. Environment drift, stale incidents, and grader changes become product-like responsibilities with owners.

When deadlines collide with evidence, communicate options: retain until deprecation, narrow supported segments, disable effects, use manual fallback, accept an explicitly authorized bounded risk, or retire. The Agentic AI Engineer should not convert organizational urgency into fabricated compatibility.

The benefit of disciplined change control is not bureaucracy for its own sake. It is the ability to move quickly where evidence is strong, stop precisely where it is weak, and know which version owns every state and effect during the transition.

Same final score, worse actions

Old and candidate both complete 88 of 100 fixtures. Candidate makes twice as many inventory reads and ignores one stop. Outcome parity hides a contract regression. Disposition is reject or new claim, not compatible.

Higher score, invalid state

Candidate fills missing fields successfully but writes unverified remote output to authoritative memory. Future runs inherit corruption. The state layer blocks release.

Fewer tokens, missing telemetry

Candidate appears cheaper after an observability adapter drops cached-input and retry events. Authoritative call counters disagree. Repair measurement before cost comparison.

Newer protocol, broader authority

A new remote card advertises reservation support and the client exposes it dynamically. Protocol adoption created permission. Quarantine; local contract review is required.

Successful migration, impossible rollback

Candidate reads old checkpoints and writes a new schema. The team calls it backward compatible but the old runtime cannot resume after one step. The cohort lacks rollback after point of no return. Drain, add reverse transform, or use roll-forward with explicit authority.

Patch-only verification

A failing fixture passes after code change. No full replay occurs. The patch might damage another segment or authority path. Run the matched suite and relevant adversarial set.

Several changes, confident cause

Model, prompt, tools, and grader all change. Completion rises. The report credits the model. Bundle evidence cannot isolate the cause. Stage isolated experiments or state attribution uncertainty.

Old claim left current

Candidate becomes default, but documentation cites replay results for the old BVS. Operators assume evidence applies. Version every claim and expire the predecessor.

Canary without state inventory

Candidate is stopped, yet its in-flight runs continue under untracked checkpoints and remote tasks. Exposure is not contained. Inventory and version ownership are release prerequisites.

Rollback as effect reversal

Deploying the old code does not undo a reservation created by the candidate. Reconcile, compensate, or transfer under domain authority.

Exercises

Exercise 1: construct a BVS

Given a model ID, prompt repository, five tools, policy engine, context builder, workflow runtime, MCP adapter, evaluator, and tracing configuration, produce a complete behavior manifest. Identify hidden versions and owners.

Five points require model/instruction/tool/policy/context/runtime/protocol/state/evaluation/telemetry coverage. Missing evaluator or state schema caps the inventory score.

Exercise 2: preregister replay

Write a plan for a tool-description and policy change. Include old claim, hypothesis, held-out/incident/adversarial/migration sets, paired trials, invariants, segmented floors, grader calibration, exclusions, and dispositions.

Do not choose thresholds after reading candidate results. Explain which claims remain inconclusive with the available sample.

Exercise 3: analyze a trajectory matrix

Candidate final outcome improves five points; tool calls rise forty percent; two stops are delayed; effect count remains correct due to a downstream gate; p95 rises; evidence coverage improves. Choose compatible, narrowed, new claim, rejected, or inconclusive and justify each layer.

A strong answer rejects direct compatibility because mandatory stop behavior fails, credits containment separately, and proposes an isolated repair experiment.

Exercise 4: migrate two checkpoints

Checkpoint A waits on an approval bound to an old proposal schema. Checkpoint B has unknown effect under an old adapter. Define transform, quarantine, revalidation, reconciliation, ownership, rollback, and audit.

Automatic failure for carrying approval to changed semantics or retrying the unknown effect.

Exercise 5: disposition protocol drift

Assess a stable dated tool spec, experimental task utility, release candidate, A2A pinned tag, advancing main branch, changed Agent Card, and unchanged schema with changed cancellation. State what is normative, volatile, pinned, tested, or insufficient.

Exercise 6: retire safely

Create a retirement checklist for runtime, queues, state, remote tasks, credentials, adapters, dashboards, documentation, and claims. Include late callbacks and the old reconciliation worker.

Exercise 7: hostile review

Argue against your preferred candidate. Find a hidden harness change, survivor bias, a missing consequence segment, a false rollback assumption, an old state class, and an external authority not consulted. Design one discriminating test per objection.

Complete AR-13 v1.0.0

The dossier now contains BVS B19-A and candidates, frozen environment and harness identities, claim/hypothesis, replay set and results, layered trajectory/effect diff, state migration matrix, protocol/tool compatibility matrix, cohort gates, stop evidence, retirement inventory, and disposition.

The planted action-policy/v2 is rejected despite higher completion because it makes extra reads and delays mandatory stops. Repaired v2.1 remains a bounded candidate subject to matched replay and approval-only cohort. The inventory schema change is compatible only for explicitly unit-bearing item counts. The changed remote cancellation behavior narrows the A2A claim to read-only low-consequence use. Old ambiguous artifacts quarantine.

AR-13 v1.0.0 does not say FieldOps is enterprise-ready or future-proof. It says which exact version set preserved which contract under which tests, what changed, what remains unknown, and how old evidence was retired. Replay cannot cover future inputs or remove remote opacity.

Chapter 20 receives a versioned portfolio rather than a bag of latest components. Leadership must now decide which mechanisms to retain, reduce, repair, reuse, platformize, or retire based on repeated evidence and a named receiving owner.

Evidence and authority boundary

CLM-037 is supported by evaluation work that examines repeated trials, trajectories, state, harness validity, and observability versions. CLM-038 is a normative synthesis using canary comparison, progress/checkpoint practice, protocol versioning, and state migration reasoning. No cited source proves arbitrary component compatibility.

Provider model IDs, protocol states, SDK behavior, telemetry attributes, graders, quotas, and prices are volatile. Reverify them at release. MCP experimental and RC material remains labeled as such; A2A specifications and protobuf are pinned. Vendor cases and benchmarks do not become current rankings.

The Agentic AI Engineer owns the BVS, local replay, migration contracts, compatibility evidence, and agent-specific rollout proposal. Evaluation specialists review methodology and uncertainty. Platform, SRE, security, IAM, privacy, procurement, enterprise architecture, domain, product, finance, and incident authorities retain their decisions. Humans keep approval, stop, takeover, risk acceptance, and retirement authority.

A change is ready only when its evidence and ownership are ready. A version number can name a candidate. It cannot earn the claim.