Name Errors and Judgment Methods
Turn raw failures into a consequence-aware taxonomy and calibrate deterministic, model, human, domain, and authority judgments without hiding disagreement.
“Wrong” is not an engineering diagnosis
Chapter 13 froze MD-06 cases with population, provenance, slices, split integrity, and known gaps. The records describe expected behavior and source state. Mosaic Desk now needs to classify failures consistently enough to compare system changes.
Consider three proposals:
- valid schema, current citation, but the qualifying condition is omitted;
- correct abstention because no eligible evidence exists;
- concise supported answer judged below a longer answer that repeats irrelevant detail.
Calling all three “bad” destroys the information needed to improve the system. The first is an unsupported or overgeneralized claim with a consequence. The second is expected bounded behavior. The third may be an evaluator bias rather than a product failure.
An error taxonomy names what happened. A judgment method states who or what can credibly decide it. These are separate artifacts.
Start from consequence, not grammar
For each behavior clause, ask what failure could occur and why it matters.
Mosaic’s initial taxonomy includes:
- state error: answer, partial, abstain, escalate, or fail-closed state is wrong;
- unsupported claim: proposal content exceeds included evidence;
- qualifier loss: condition, negation, exception, date, or scope disappears;
- wrong evidence: citation resolves but does not support the claim;
- fabricated provenance: source handle, revision, or span does not exist;
- stale-source use: superseded evidence supports a current claim;
- authorization breach: forbidden content enters candidates, context, output, or trace;
- conflict erasure: disagreement is hidden or one source is chosen without authority;
- inappropriate abstention: answerable bounded request is refused;
- missing abstention: unsupported request receives a confident proposal;
- schema/interface failure: typed result violates structural or semantic contract;
- instruction-boundary failure: untrusted data changes control behavior;
- prohibited effect claim: text asserts approval, contact, scheduling, order, or mutation;
- language/meaning failure: material intent or support changes in a declared language segment;
- style or usability defect: relevant information is obscured without changing correctness;
- resource failure: time, token, capacity, or cost budget is exceeded.
The list is not universal. Merge or split categories only when the distinction changes detection, consequence, control, or disposition.
Characterize one specimen across dimensions
An error record needs more than a category.
criterion: the behavior clause or measurement rule violated;category: what failed;consequence: the possible effect on reviewer, customer, system, or evidence;severity: locally defined consequence band;detectability: whether deterministic checks, routine review, specialist review, or incident evidence can reveal it;segment: language, appliance, evidence state, risk, and other relevant slices;responsibleLayer: current hypothesis from the Chapter 12 trace;control: prevention, detection, containment, or recovery mechanism;disposition: accept as correct, repair, abstain, escalate, fail closed, block, or investigate;judgeMethod: instrument credible for the field;uncertainty: disagreement, missing evidence, or unresolved authority;owner: who improves the mechanism and who makes the formal decision.
Criteria, error category, consequence, severity, detectability, segment, and disposition are distinct evaluation fields. Holistic evaluation, provider criteria guidance, and NIST’s pilot measurement work motivate structured evaluation, while this book’s taxonomy remains task- and risk-specific rather than a standard. [CLM-040]
Severity cannot be inferred from model confidence, judge score, or frequency. A rare another-tenant disclosure can be more consequential than many tone defects. Detectability is also independent: a severe fabricated citation may be easy to catch deterministically, while a subtle domain qualifier loss may require expert review.

CLM-040; it defines no universal severity.Long description
A central three-dimensional error specimen card is connected to seven distinct evidence cards. Consequence and severity are separate, detectability points to a searchlight, segment to a mosaic tile, control to a barrier, and disposition to a routing gate. An owner badge sits outside the technical record.
Match each claim to the cheapest credible judge
“Cheapest” alone encourages weak evidence. “Most expert” alone makes evaluation impossible to scale. Choose the least costly method that is credible for the claim and consequence.
Deterministic checks
Use software for schema validity, allowed enums, source-handle existence, revision match, permission, exact prohibited effects, budget accounting, and trace completeness. Deterministic checks are reproducible but only judge encoded rules.
Reference or rule-based comparison
Use references for exact identifiers, known state transitions, bounded spans, and fields with stable equivalence rules. A single reference string is weak for open-ended language and can penalize valid variants.
Model graders
Use a versioned model grader for bounded language criteria such as whether a claim is supported by a supplied span, after calibration. Record model, access path, prompt/rubric, order, decoding, schema, retries, and unavailable internals. A grader score is an instrument reading.
Trained human review
Use trained reviewers for ambiguity, usability, nuanced support, and calibration anchors. Record training, rubric, blinding, order randomization, time, disagreement, and adjudication. Human judgment is not automatically consistent or domain-authoritative.
Domain specialist
Use a qualified domain specialist for source correctness, policy meaning, appliance applicability, and consequential ambiguity. The specialist may determine what evidence supports, not whether the product should be released.
Accountable authority
Product, domain, legal/privacy, security/safety, or release owners make formal decisions within documented authority. Evaluation evidence informs but does not replace them.
Calibrate the instrument before scaling it
Build a calibration set distinct from the final holdout. Include clear pass/fail cases, close calls, expected abstentions, conflict, different languages, short and long answers, citation-strength differences, and deliberate order/style traps.
For each criterion:
- define the rubric and evidence available to the judge;
- collect independent credible human or domain labels;
- preserve disagreement rather than forcing easy consensus;
- run deterministic and model methods blind to the anchor;
- compare agreement by category, consequence, and segment;
- inspect false passes and false failures;
- reverse pair order and control verbosity/style;
- revise the instrument without tuning on protected holdout;
- freeze version and escalation rules;
- state unsupported uses.
Agreement is not validity by itself. Two graders can agree because they share a blind spot. A grader can also disagree with a human because the rubric is ambiguous or the human label is weak. Inspect specimens.
Test position, verbosity, and self-preference
Pairwise model judging is attractive because it avoids inventing an absolute scale. It still has systematic failure surfaces.
Create content-equivalent pairs:
- swap answer A and B while keeping text fixed;
- add irrelevant but polished detail to one answer;
- change formatting and headings without changing meaning;
- compare output from the judge’s own model family with another path;
- insert a correct concise abstention against a longer unsupported proposal.
LLM judges can exhibit position, verbosity, and self-preference biases. MT-Bench/Chatbot Arena research and reward-model benchmark evidence document such concerns under their judge, task, and data setups. Bias direction and magnitude vary; human preference is not universal correctness and a reward-model leaderboard is not a Mosaic release decision. [CLM-041]
LLME-CASE-005 is the bounded case: its study reported useful agreement in a specific setup while documenting biases and limitations. Mosaic borrows the calibration requirement, not an agreement rate or judge choice.
Preserve disagreement as evidence
Use a disagreement log with:
- case, criterion, and evidence snapshot;
- deterministic result if applicable;
- randomized model-grader results and versions;
- independent human/domain judgments;
- order and presentation variants;
- disagreement category: rubric ambiguity, missing evidence, grader bias, reviewer inconsistency, or domain dispute;
- adjudicator and authority needed;
- resolution, unresolved status, or case exclusion;
- follow-up to rubric, case, control, or system.
Model graders should be calibrated against credible human or domain judgments and retain disagreement/limitations. Official judge-evaluation and grader documentation plus model-grading research support calibration, while human and domain judgments can also be inconsistent. Calibration measures an instrument for a bounded criterion; it does not confer ground-truth or high-stakes decision authority. [CLM-042]
Separate measurement uncertainty from system uncertainty. If graders disagree about a fixed proposal, that is measurement uncertainty. If the proposal correctly reports conflicting sources, that is system/domain uncertainty. One cannot be repaired by increasing the other’s confidence score.

CLM-041 and CLM-042; it does not rank human or model infallibility.Long description
A bright three-dimensional staircase begins with a code-check terminal, then a reference card, model-grader cube, trained-human desk, domain-specialist station, and formal authority gate. Cases travel only as high as their ambiguity and consequence require. Disagreement cards remain attached between levels.
Build a rater guide that constrains interpretation
For each criterion, specify:
- plain-language definition;
- evidence the rater may use;
- inclusions, exclusions, and boundary cases;
- positive, negative, and abstention examples;
- distinction from adjacent categories;
- consequence and escalation conditions;
- whether partial credit exists;
- required rationale and citation;
- what the rater must not decide.
Train on cases outside the final measurement set. Measure agreement by criterion and segment, not just overall. Review low-agreement items for rubric defects or genuine ambiguity. Do not discard difficult cases solely to raise agreement; retain them in a disputed or specialist queue when they represent intended use.
Require evidence-shaped judge outputs
A grader should not return only a number. Use a typed record containing:
- case, criterion, rubric, grader, and evidence-snapshot versions;
- verdict from a bounded enum;
- evidence spans or deterministic facts used;
- short rationale tied to the rubric;
- uncertainty or cannot-judge state;
- detected order/presentation condition;
- escalation target when the criterion exceeds the instrument;
- no authority or effect field.
Validate the record like any other generated output. A schema-valid grader rationale can still cite the wrong span or apply the wrong criterion. Preserve raw grader output, normalized record, validator errors, retry state, and terminal disposition.
Use cannot-judge deliberately. A general grader presented with a domain policy dispute should stop rather than invent authority. A missing span should yield insufficient evidence, not a low correctness score. Bounded failure states improve measurement because they separate unsupported judgment from negative judgment.
Interpret calibration by failure mode
An overall agreement percentage can hide asymmetric risk. Break calibration into false passes, false failures, cannot-judge use, order flips, verbosity flips, segment disagreements, and critical-case disagreements.
A false pass on a prohibited effect matters differently from a false failure on tone. A grader that agrees overall but misses every Gujarati qualifier is not credible for that segment. A grader that refuses ambiguous domain cases may be useful as a triage instrument when the escalation route is designed.
Do not tune a threshold until the score meaning is stable enough to support the criterion. If a tiny score change flips many cases, inspect raw rationales and distribution. Calibration drift can come from grader-model updates, prompt changes, source changes, rater-guide revisions, or population shifts. Any of these triggers revalidation.
Retain a small bias battery outside ordinary calibration: order swaps, length controls, self-family outputs, equivalent paraphrases, correct abstentions, adversarially polished unsupported claims, and multilingual pairs. Run it whenever the grader identity changes.
Route error dispositions separately from product priorities
The taxonomy can recommend a technical disposition: fix the validator, revise context packing, improve labels, block the candidate, or escalate a source dispute. It cannot decide the organization’s acceptable residual risk or product tradeoff.
Record both owners. The LLM engineer may own a grader calibration failure. A domain owner may own an ambiguous policy label. Privacy/security owners may own unauthorized-data disposition. Product and release authorities decide whether evidence supports a scoped progression. Keeping these fields separate prevents the person who built the metric from silently becoming the decision authority.
Keep correct abstention out of the error bucket
The taxonomy must represent expected bounded behavior. If source state is no-evidence, an abstaining proposal can pass even though task completion is limited. If evidence is available and low consequence, unnecessary abstention may be a usability or state error. If sources conflict, escalation may be correct.
Judge state first, then content within the state. A rubric that always rewards detailed answers will punish safe abstention and encourage unsupported prose. Include explicit no-evidence and out-of-authority cases in calibration.
Preserve provider and judge independence
Managed and open-weight candidate systems share the same criteria and cases. Their judge adapters may differ, but comparisons must record coupling.
A managed grader can change behind an endpoint and hide internal versions. Recheck documented behavior, retain sampled outputs, and mark unavailable identity. An open-weight grader offers artifact and runtime control but adds tokenizer, template, quantization, serving, and hardware responsibilities. Neither path is inherently independent.
Where possible, avoid using the candidate to judge itself. Where impossible, label the coupling and compare against external human/domain anchors. Do not upload protected cases to a provider without authorized data terms.
Failure injection: first and longest wins
Construct two synthetic proposals with equivalent supported content. Proposal A is concise. Proposal B is longer and repeats background details. Run a pairwise grader four ways: A/B, B/A, equalized formatting, and shortened B.
The fixture’s deliberately biased grader selects the first answer, then uses length as a tie-breaker. The apparent winner changes under order reversal. A deterministic support checker finds both equivalent on the bounded criterion. Trained reviewers record no material support difference.
Disposition:
- fail the model grader for preference use on this criterion/version;
- preserve every raw judgment and order;
- route support to deterministic/span checks plus calibrated review;
- revise the rubric and grader prompt on calibration cases only;
- retain an order-reversal regression test;
- grant no source, warranty, safety, or release authority to the grader.
This synthetic result proves the bias fixture works, not that a named commercial or open-weight model has the same bias rate.
Practice: calibrate ten failures
Take ten Chapter 12/13 specimens. For each:
- name criterion, category, consequence, severity, detectability, segment, control, disposition, and owners;
- select the cheapest credible judge and explain why weaker methods are insufficient;
- write a rater-guide entry;
- collect blinded independent anchor judgments;
- randomize order and control length/style for model grading;
- compare by category and segment;
- preserve disagreements and limitations;
- route domain correctness to specialists;
- keep correct abstention visible;
- freeze taxonomy, judge, rubric, and escalation versions.
Pass when another reviewer can reproduce each judgment route, see bias tests and disagreement, and distinguish technical measurement from formal authority.
The completed MD-06 evidence system
Mosaic exits Chapter 14 with a consequence-aware error taxonomy, severity/detectability matrix, criterion-to-judge map, calibration set, randomized model-grader traces, trained-rater guide, domain escalation rules, disagreement log, correct-abstention treatment, and frozen evaluator versions. MD-06 is now complete enough for controlled comparison, not release. Chapter 15 will change one system variable family at a time and preserve segment regressions, uncertainty, and negative findings in MD-07.