NEWProduction Web Themes & Turnkey ArchitecturesGet Lifetime Pass ($199) →
KNKomal Nakrani
Get All Access
ThemesDocsAll-Access PassGet All Access ($199)
Book overview
05/Applied AI Engineering

Make Task Data Representative

Build a permitted, versioned task-data substrate whose population, provenance, labels, segments, leakage, missingness, and limitations can support or reduce the behavior claim.

A valid row can still be unfit evidence

Patchwork’s first synthetic catalog record passes its schema. It has a listing ID, title, category, seller label, dimensions, unit, freshness date, permission state, and image-feature vector. Every required property exists. The validator returns green.

The row is still wrong for the behavior claim.

Its diameter field describes outer housing diameter, while the task requires shaft diameter. The seller label “fits most garden pumps” was copied into a compatibility field. A near-duplicate of the listing appears in both development and release evaluation. The image vector was constructed from the same class template used to assign the expected category. A test can achieve an excellent score by learning the generator rather than the task.

Schema validity answers whether a record has the expected structure. It does not answer whether fields mean what the product thinks, whether the population resembles the supported task, whether labels came from competent judgment, whether processing leaked the answer, or whether the organization may use the data for this purpose.

Chapter 4 selected a mechanism portfolio conditional on data. This chapter executes that condition. If Patchwork cannot build evidence for its required attributes and segments, the product claim must narrow or stop. A stronger model cannot repair an undefined label or unauthorized source.

Representative of what?

Data is never representative in the abstract. It is representative enough only relative to a specified:

  • task: the decision or behavior under evaluation;
  • population: eligible actors, items, contexts, and affected groups;
  • time: collection and use period, including known changes;
  • segment: conditions under which behavior or consequence can differ;
  • consequence: errors and exclusions the evidence must reveal;
  • claim: the exact statement the team wants to support.

For Patchwork, the current task is evidence-backed candidate discovery for selected low-consequence pump-seal families when the buyer can supply a family and required measurements. The data does not represent all repair parts, all buyers, all sellers, all images, or all marketplaces. The local companion uses fictional seeded records. It represents designed software states, not a real distribution.

This distinction changes sampling. A random draw from whatever catalog rows are easiest may match overall frequency and still omit unitless legacy listings, sparse categories, conflicting attributes, inaccessible images, or excluded high-consequence cases. A challenge set that deliberately includes those states can test controls but cannot estimate prevalence.

The data card must name whether a set supports development mechanics, challenge coverage, evaluator calibration, release evidence, or production monitoring. One set rarely supports every purpose.

The unit of observation

Before counting records, define what one record represents.

A listing row represents one version of seller-provided and system-derived catalog evidence. A query instance represents a user’s stated need and available inputs at a time. A candidate judgment connects a query version to a listing version under criteria and reviewer provenance. A task instance includes the interaction state and eventual next decision.

Collapsing these units creates leakage and false claims. If a listing appears under several seller IDs, a row-level split can place near-identical content in development and evaluation. If several queries come from one task session, a query-level random split can expose the same intent on both sides. If labels are updated after a return, time can leak future evidence into an earlier evaluation.

Patchwork therefore records stable group keys:

  • canonical product family;
  • normalized seller/listing cluster;
  • synthetic generation template;
  • task-session family;
  • source and timestamp;
  • domain-rule version.

Splits operate at the group boundary appropriate to the claim, not merely by row.

Data documentation is a question system

Dataset documentation approaches recommend recording motivation, composition, collection, processing, intended use, and limitations. A well-designed card creates accountability questions and makes missing facts visible. It does not prove permission, fairness, quality, or fitness. [CLM-021]

Patchwork’s data card contains the following minimum sections.

Identity and version

Name, semantic version, immutable content hash, generator/code version, seed, creation date, owner, review status, and superseded versions.

Motivation and claim

Which behavior clauses and mechanism decisions need this data? Which decisions must not use it?

Unit and population

Define record types, eligible population, excluded population, time window, unknown population, and affected groups.

Composition

Counts by declared segment, missingness, duplicate cluster, source type, label state, permission, freshness, and split. Counts describe this dataset only.

Provenance and permission

For each field and artifact, record source, owner, collection method, license or organizational permission, user expectation, purpose, access, retention, correction, and deletion. A card cannot grant permission; it records the decision and authority.

Label and judgment process

Define criteria, source evidence, reviewer competence, instructions, disagreement, adjudication, uncertainty, and version. Use label or judgment, not unqualified ground truth.

Processing

Record normalization, unit conversion, deduplication, feature generation, redaction, augmentation, transformations, and their versions.

Splits and leakage

Name development, tuning, evaluator-development, calibration, challenge, and release-evaluation partitions; grouping rules; freeze dates; contamination checks; and known shared dependencies.

Known skews and gaps

List underrepresented categories, missing attributes, seller-source limitations, image ambiguity, historical bias, unknown intersections, and excluded use.

Prohibited uses

State claims and actions the dataset cannot support. Synthetic Patchwork data cannot establish real-world performance, prevalence, fairness, buyer intent, seller behavior, compatibility safety, or release readiness.

Refresh and retirement

Define schema, catalog, policy, population, time, label, model, and incident triggers. State who decides whether to rebuild, append, freeze, or retire.

A documentation field without evidence remains unknown; it does not become true because the template is complete.

Population, sample, label, and segment

Four objects should remain visible.

  1. Intended task population: all instances to which the product would apply under the contract.
  2. Observed or constructed sample: records actually available for development or evaluation.
  3. Judgment process: how expected behavior or relevance was assigned.
  4. Analysis segments: predeclared slices where evidence or consequence may differ.
Nested view of intended task population, observed sample, label process, and intersecting analysis segments, with unknown and excluded groups kept visible.
F05.1 - Population, sample, label, and segment composition. A valid sample can still leave unsupported groups, unknown populations, and consequential intersections outside its evidence.

The sample does not inherit the population by naming it. A dataset containing many common 16 mm listings may not represent rare seal families or unitless legacy rows. A label process based on seller claims may represent seller wording, not domain compatibility. A photo set with clear centered images may not represent worn, occluded, low-light, or wrong-object photos.

Research evaluating demographic effects and intersectional disparities in specific face-analysis systems demonstrates why aggregate composition and performance can hide material groups. Those studies concern particular technologies, datasets, labels, and historical periods. They are not evidence about Patchwork or every vision product. The transferable discipline is to predeclare consequence-relevant segments and inspect intersections rather than let an aggregate erase them. [CLM-024]

For Patchwork, segments arise from the contract and failure model, not protected-attribute fishing. They include:

  • explicit versus missing unit;
  • complete versus incomplete required dimensions;
  • common versus long-tail catalog family;
  • current versus legacy schema;
  • single versus conflicting source;
  • fresh versus stale listing;
  • authorized versus unauthorized source;
  • clear versus ambiguous synthetic image feature;
  • unique versus duplicate listing cluster;
  • literal identifier versus descriptive query;
  • included versus prohibited consequence category.

Where demographic or accessibility evidence becomes relevant to user burden, it requires purpose, authority, careful terminology, and appropriate methods. The synthetic dataset does not fabricate demographic fields merely to appear responsible.

Provenance is field-level

A dataset-level source statement is too coarse. One listing can combine seller text, catalog-team normalization, inferred category, image feature, buyer correction, and domain review. Each field has different authority.

Record a provenance tuple:

source_id, source_version, field, method, actor_or_system, timestamp, permission, confidence, transformation, review, retention.

Patchwork’s synthetic rows distinguish:

  • seller_stated: fictional seller-provided value;
  • domain_normalized: transformed under a versioned constructed rule;
  • system_inferred: produced by a mechanism and never displayed as seller fact;
  • user_entered: supplied for the current task;
  • synthetic_fixture: designed for local testing only.

The distinction lets the product say, “Listing states 16 mm” rather than “This part is 16 mm” when the source has not been verified. It also prevents a model inference from being written back as truth.

Provenance supports correction. If a domain rule changes, the team can identify which normalized values and judgments depend on it. If a user disputes a listing, support can locate the source rather than treating the model output as origin.

Permission is not a boolean copied forever

allowed: true hides purpose, scope, actor, retention, and change.

For every data source, record:

  • permitted purpose and behavior clauses;
  • allowed users, services, and environments;
  • source and destination boundaries;
  • retention and deletion behavior;
  • whether derived features may persist;
  • whether evaluation, model development, monitoring, or support use differs;
  • authority and decision record;
  • expiry or review trigger.

Privacy frameworks and adversarial knowledge bases help teams ask lifecycle and threat questions, but they do not grant legal permission or supply a complete application threat model. [CLM-026]

Patchwork’s local companion uses only fictional data created for the book. That removes real personal-data exposure from default execution. It does not prove that a future real dataset may be used. A production team must make purpose and authority decisions independently.

Image data deserves separate care. A buyer photo can contain people, location cues, labels, or unrelated objects. A feature vector can retain information and requires its own purpose and deletion behavior. “We only store embeddings” is not automatically minimization.

Labels are decisions with lineage

A label is an output of criteria, evidence, people or systems, and time. Call it ground truth only when the domain supports that claim and the authority agrees.

Patchwork needs different judgments:

  • candidate relevance for discovery;
  • explicit attribute match;
  • deterministic incompatibility;
  • evidence sufficiency;
  • expected behavior state;
  • explanation faithfulness.

One label cannot collapse them. A listing may be relevant but incompatible. It may be compatible on known dimensions but insufficiently evidenced. It may be a correct candidate while the system should abstain because one critical fact is unresolved.

For every judgment, record:

  • criterion version;
  • evidence shown;
  • reviewer role and competence;
  • independent or collaborative review;
  • uncertainty and allowed states;
  • disagreement and adjudication;
  • time and source versions;
  • intended use and prohibited reuse.

Domain documentation cases show the value of recording expert involvement and label provenance. They do not imply that labels are perfect or that expertise transfers across domains.

Seller outcomes are especially dangerous labels. A purchase is not compatibility. A lack of return is not successful fit. A click is not relevance. These signals may inform a hypothesis but require a causal and error model.

Structural, semantic, and fitness validation

Use three layers.

Structural validation

Check types, required fields, ranges, allowed values, referential integrity, shape, vector length, timestamps, and version syntax. Production data-validation research demonstrates mechanisms for detecting anomalies, skew, and drift at scale. Those mechanisms do not eliminate semantic review. [CLM-022]

Semantic validation

Check whether values mean what the contract assumes. Is diameter shaft diameter? Is the unit explicit? Does category pump-seal use the same criteria across sources? Can a stale seller claim remain authoritative? Semantic rules require domain ownership and may produce unknown or conflicting states.

Fitness validation

Check whether the data can support the intended behavior claim. Does the sample cover critical segments? Are labels appropriate to the decision? Are permissions adequate? Are splits uncontaminated? Are limitations compatible with the release scope?

A record can pass structure and fail semantics. A dataset can pass structure and semantics for included rows and fail fitness because the sample excludes consequential populations.

The validation report keeps all three dispositions. Do not summarize them as one “data quality score.”

Missingness is information

Missing values arise for different reasons:

  • not collected;
  • not applicable;
  • unavailable from source;
  • redacted;
  • collection failed;
  • legacy schema omitted it;
  • value conflicts;
  • user declined;
  • permission forbids use;
  • not yet reviewed.

Encoding all as null erases the decision. Patchwork uses an evidence state with status, reason, source, and observed_at. A missing critical unit triggers clarification or abstention. An unauthorized value behaves differently: the system must not use it even if technically present.

Missingness can be segment-dependent. Sellers with older tools may omit structured attributes. Long-tail categories may have less domain review. A model that performs well on complete records can create unequal coverage. Record coverage and burden, not only accepted-case quality.

Do not impute a value merely to satisfy a schema. Imputation is a mechanism with assumptions, evidence, and user-facing consequences. For Patchwork’s critical dimensions, current v0.1 prohibits inference when alternative values change compatibility.

Duplicates and correlated records

Exact row equality is the easiest duplicate. Near-duplicates are more common:

  • same listing copied by resellers;
  • title punctuation or image crop changes;
  • one canonical part under multiple seller IDs;
  • revised listing versions;
  • several queries generated from one template;
  • multiple judgments from the same evidence packet.

If related records cross development and evaluation, the system can appear to generalize while recognizing content or generator patterns. Deduplicate or group before splitting. Preserve a cluster ID and the matching method, including limitations.

Patchwork’s synthetic generator intentionally creates L-004 and L-004-COPY with different seller labels but the same canonical item and near-identical features. The first naive row split places one in development and one in release evaluation. The leakage detector must fail. The corrected split groups them by canonical_cluster_id.

Leakage is any path by which held-out evidence influences development

Leakage is broader than the target column appearing in features.

  • the same item or near-duplicate crosses partitions;
  • evaluation labels influence feature or prompt design;
  • a global normalization statistic uses held-out data;
  • the generator encodes class directly in a synthetic feature;
  • a reviewer remembers evaluation cases while tuning criteria;
  • a model provider trained on supposedly held-out public examples;
  • future catalog state appears in a historical decision;
  • evaluator development uses release cases;
  • a retrieval index contains evaluation-only answers.
Leakage path map separating source population, development, tuning, evaluator development, challenge, and frozen release evaluation, with prohibited duplicate, time, label, feature, and reviewer paths.
F05.2 - The leakage path map. Release evidence remains credible only when content, labels, features, time, reviewers, and evaluator development cannot carry answers across the freeze.

Use separate partitions for distinct decisions:

  • development: build transformations and debug behavior;
  • tuning: choose parameters or prompts;
  • evaluator development: build and calibrate judgment methods;
  • challenge: exercise known rare and adversarial states, without prevalence claims;
  • release evaluation: frozen cases used for disposition;
  • monitoring reference: production comparison where authorized.

The release set should remain inaccessible to ordinary iteration. When it is used, record the decision. Repeated use converts it into development evidence and requires a new release set.

Data cascades are organizational

An ambiguous field does not remain a local data issue. It shapes sampling, labels, features, model behavior, evaluation, interface claims, and operations. Research on data cascades describes how neglected data work and unclear responsibility can compound across sociotechnical stages. The qualitative findings have sampled-context limits; the diagnostic principle remains useful. [CLM-023]

Patchwork’s diameter ambiguity can cascade:

  1. a seller form asks for diameter without definition;
  2. ingestion accepts values without units;
  3. normalization assumes millimeters;
  4. training and evaluation treat the field as comparable;
  5. the ranker rewards apparent closeness;
  6. generated text calls it a match;
  7. the interface shows compatibility;
  8. returns become a noisy label;
  9. later models learn from the contaminated outcome.

The repair begins upstream: define the field, preserve unknown states, correct the form and migration, assign ownership, and replay affected evidence. Model tuning alone deepens the cascade.

Synthetic data is a test instrument

Synthetic data is valuable for this book because it is safe, deterministic, distributable, and capable of injecting failure states. It can prove that code detects a duplicate, blocks an unauthorized record, handles missing evidence, and reproduces a split.

It cannot prove:

  • real population frequency;
  • real user language or behavior;
  • seller incentives or error patterns;
  • real image variability;
  • product quality or fairness;
  • compatibility outcomes;
  • production latency or cost;
  • release readiness.

Patchwork’s vectors are deliberately constructed. A successful vector-like retrieval test demonstrates interface and ranking mechanics only. Multimodal research and product cases show feasible approaches but remain data-dependent; their results cannot be imported into the synthetic case. [CLM-025]

The data card prints case_status: fictional-synthetic and representativeness_claim: none. Tests fail if these fields disappear.

Build a sampling and segment plan

A sample should be designed from decisions backward. Start with the claim and ask which conditions could make it false.

Define the sampling frame

List the sources from which records could be drawn, their coverage, and the population each source misses. A catalog export may exclude deleted listings. Support cases overrepresent failures that reached support. Purchase events exclude abandoned searches. Manual review queues overrepresent uncertain cases. No source is “the population” by default.

Patchwork’s synthetic frame is explicit: a generator produces constructed listings and queries from named templates. The frame cannot include real behavior. Its only valid population is the set of states the generator can emit under version 0.1.0.

Predeclare material segments

Segments arise from behavior clauses, mechanism assumptions, and consequence. Record why each segment matters and which decision it can change. Avoid slicing every available field after results arrive; that creates unstable stories and multiplicity without a decision model.

For Patchwork, missing unit matters because it changes PF-B06. Unauthorized source matters because it changes inclusion regardless of relevance. Long-tail category matters because representation and rank signals may be sparse. Duplicate cluster matters because it can invalidate release evidence.

Decide coverage purpose

Use different samples for:

  • frequency estimation, where a defined sampling frame and weights may support prevalence within limits;
  • contract coverage, where every clause and state needs representative examples;
  • challenge testing, where rare or constructed failures are deliberately overrepresented;
  • diagnosis, where cases concentrate around an observed regression;
  • calibration, where judgment distributions and score behavior matter;
  • release disposition, where a frozen set supports a bounded decision.

Do not merge their counts. A challenge set with 30 percent unauthorized sources says nothing about production prevalence.

Represent intersections deliberately

Single dimensions can look healthy while intersections fail. A long-tail family with missing dimensions and legacy units may behave differently from each factor alone. Predeclare intersections where mechanism or consequence plausibly changes, then record sample insufficiency rather than inventing confidence.

The number of intersections grows quickly. Prioritize through consequence, expected exposure, mechanism sensitivity, and affected-group burden. Unknown is a valid result.

Preserve excluded and unknown populations

A scope table should show:

Population Status Reason Product behavior
Supported low-consequence family with required evidence included for evaluation contract scope eligible for candidate assistance
Missing critical measurement observed state, not accepted response population insufficient evidence clarify or abstain
Legacy unitless listing critical challenge segment semantic ambiguity exclude from assisted compatibility
High-consequence category prohibited authority and consequence boundary ordinary/non-assisted route only
Unknown catalog family unknown population no evidence out of scope

This prevents the release report from silently treating excluded cases as successes or erasing them from coverage.

Design a label study rather than a labeling task

Before scaling annotation, test whether the criterion can be applied consistently and usefully.

Write criterion cards

For each judgment, define purpose, inputs, allowed outputs, ordinary examples, boundary examples, counterexamples, uncertainty, escalation, prohibited inference, and downstream use. A relevance card must say whether a visually similar seal is relevant when a critical dimension conflicts. A compatibility card may be prohibited entirely without authoritative domain evidence.

Select reviewers by competence

Language familiarity may be enough for lexical relevance. Domain evidence may require a catalog specialist. Policy or safety disposition requires another authority. Reviewer identity is not merely demographic metadata; role, training, access, and decision right determine what the judgment can support.

Preserve independent judgments

Early independent review reveals criterion ambiguity. A consensus meeting can hide disagreement by producing one label. Keep the initial decisions, rationales, and uncertainty before adjudication.

Diagnose disagreement

Disagreement can reveal:

  • unclear criterion;
  • insufficient evidence;
  • different domain assumptions;
  • ambiguous source semantics;
  • reviewer error;
  • legitimately variable acceptable behavior;
  • an authority dispute.

Do not automatically resolve it by majority vote. If experts disagree because the catalog lacks a unit, the correct label may be unresolved, not whatever two of three selected.

Calibrate the process

Use shared cases to refine instructions and measure agreement appropriate to the label type. Agreement is not validity. High agreement can reflect a shared wrong assumption. Compare against source evidence, adjudication, and downstream consequence.

Version and limit reuse

A relevance judgment created for candidate retrieval may not support compatibility, explanation faithfulness, or user value. Store criterion_id, version, evidence packet, reviewer role, outcome, uncertainty, and intended use. When criteria change, do not overwrite old labels.

Patchwork’s companion uses constructed expected states, not human annotations. Tests can verify that the program follows the fixture. They cannot establish that the fixture represents domain truth.

Rebuild a contaminated split

The first Patchwork split uses a hash of listing_id. It looks deterministic and balanced. It is invalid because two near-duplicate seller listings have different IDs. One enters development and one enters release evaluation.

Step 1: freeze the contaminated evidence

Record the dataset hash, split algorithm, seed, duplicate detector version, and report. Do not delete the failed split; it demonstrates why earlier results cannot support release.

Step 2: identify relationship keys

Construct canonical_cluster_id from authorized catalog relationships and deterministic synthetic generator metadata. Add generation_template_id and task_family_id. A production system would need careful review because automated duplicate clustering can merge distinct items or miss copied content.

Step 3: split groups, not rows

Assign whole canonical clusters and generation templates to one partition. If the same template produces both development and release cases, the release set may measure template variation rather than generalization. Hold out entire templates for selected claims.

Step 4: fit transformations on development only

Vocabulary, normalization statistics, feature scaling, deduplication thresholds, rank weights, prompt examples, and evaluator criteria must not learn from release evidence. Apply frozen transformations to held-out records.

Step 5: re-run contamination checks

Check exact IDs, canonical groups, near-duplicate text/features, source versions, timestamps, task sessions, label provenance, reviewer overlap, and generator templates. Report both detected and undetectable routes.

Step 6: inspect segment composition

A clean split can be unusable if critical segments disappear. Rebalance at the group level or state that the release set cannot support those segments. Never move one duplicate row simply to improve counts.

Step 7: seal release evaluation

Restrict ordinary development access, record the content hash, and log each use. After repeated release decisions or criterion changes, retire it and create a new set from an uncontaminated process.

This process reduces known leakage. It does not prove independence from public pretraining, reviewer memory, or unknown source relationships. State those limitations.

Separate development evidence from release evidence

Developers need fast, inspectable cases. Release authorities need evidence that was not optimized directly. Evaluator designers need separate cases again, especially when building model-based or heuristic graders.

Use an evidence ledger:

Partition May influence mechanism? May influence evaluator? May support release?
Development yes yes, if recorded no
Tuning yes, through selection yes, through selection no
Evaluator development not directly; results may expose behavior yes no
Challenge yes after discovery; then becomes development knowledge possibly no prevalence or untouched-release claim
Frozen release only after disposition, then future work treats it as used no before run yes, for scoped version and criteria

When a release case causes a fix, celebrate the finding and retire its untouched status. Add a regression case to development and replenish release evidence. Do not rerun until it passes and call the same set independent.

The ledger must also record human exposure. A reviewer who authored expected answers can still evaluate execution mechanics, but their judgment is not blinded release evidence.

Threat-model the data substrate

Task data can be attacked or can carry unsafe instructions into later components.

Source manipulation

A seller may stuff keywords, copy a competitor, falsify dimensions, or place instruction-like text in a description. Provenance identifies the source; it does not make the content trustworthy.

Label manipulation

Outcome proxies can be gamed. Purchases, clicks, and non-returns can be influenced by ranking position, promotions, friction, or fraud. Treat them as observed signals with causal limits.

Evaluation poisoning

Public or reused evaluation cases can enter development or provider training. A system may overfit the test. Preserve private or freshly constructed evidence where authority permits and report the remaining uncertainty.

Feature and image attacks

Adversarial changes can affect learned representations. Security taxonomies and adversary knowledge bases organize known patterns but do not supply likelihood or a complete Patchwork threat model. [CLM-026]

Sensitive-content propagation

Images and text may contain personal or secret information. Redaction can fail or alter task evidence. Design collection, access, retention, and deletion before using the data.

Derived-data persistence

Embeddings, labels, caches, and traces can survive source deletion. The data card must map derived artifacts and deletion/rebuild behavior.

For the local companion, every source is synthetic and no untrusted external content is loaded. Tests still inject instruction-like listing text to prove that retrieval evidence is treated as data, not as an application command.

Refresh, freeze, and retire deliberately

Data changes through append, correction, policy, schema, and population shift. A refresh is a new evidence version, not an invisible file replacement.

Trigger review when:

  • task or supported segment changes;
  • catalog schema or field meaning changes;
  • permission, retention, or source ownership changes;
  • duplicate or leakage route is discovered;
  • label criterion or reviewer process changes;
  • a consequential production failure is absent from the set;
  • missingness or segment composition moves materially;
  • model, representation, or preprocessing changes what evidence is needed;
  • the release set has been used enough to become development knowledge;
  • a source is withdrawn or corrected.

For each trigger, decide:

  • append: add new records while preserving comparability;
  • rebuild: regenerate under new semantics;
  • correct: issue a versioned erratum and trace affected judgments;
  • resplit: rebuild partitions because grouping or leakage changed;
  • retire: stop using the set for a claim;
  • fork: preserve an old contract population while creating a new one.

Never mix versions in one metric without explaining the population and semantics. Store content hashes and data-card versions with every experiment.

Resolve data disputes through the claim

“The sample is statistically large”

Size cannot repair a biased frame, missing critical segment, wrong label, or duplicate leakage. Ask which population and consequence the calculation represents.

“The schema validates, so data engineering is done”

Structure is one layer. Require semantic and fitness dispositions with named owners.

“Seller labels are the best truth available”

They may be a useful source signal. Preserve provenance and test their relationship to the required judgment. Do not silently promote marketing language to domain fact.

“Synthetic data is unbiased”

Synthetic data reflects generator choices, templates, and omissions. It can make those choices reproducible, not neutral.

“We should collect more before deciding”

Name the uncertainty and purpose. More unauthorized or semantically weak data makes the evidence problem larger.

“The data card says approved”

A card records an authority decision and its scope. It is not the authority. Verify the source record, purpose, version, and expiry.

“An aggregate looks good”

Inspect predeclared segments and intersections tied to consequence. Do not transfer findings from unrelated demographic studies as if they measured this system. [CLM-024]

More data can make the system worse

Adding records can deepen:

  • majority-category dominance;
  • duplicate leakage;
  • seller-label bias;
  • historical policy assumptions;
  • unauthorized retention;
  • attack content exposure;
  • evaluation contamination;
  • maintenance and correction cost.

Ask what uncertainty each increment reduces. If the problem is a missing long-tail segment, another million common-category rows do not help. If label semantics are wrong, more labels reproduce the error. If permission is absent, collection increases exposure.

Data quantity is not a substitute for claim discipline.

Inspect the constructed Patchwork dataset

The local generator creates a deliberately small dataset so every record can be understood. The exact counts are fixture facts, not estimates about a marketplace.

Listing examples

PW-L001 is an authorized, fresh listing in an included seal family with explicit 16 mm shaft diameter and 28 mm seat diameter. Its structured fields and title agree. It is a normal respond candidate for a matching query.

PW-L004 describes the same constructed canonical item but uses a seller marketing label. PW-L004-COPY copies the title and feature vector through another fictional seller, changes punctuation, and omits one structured field. A row-level split separates them. The group detector identifies their shared canonical cluster and generator template.

PW-L006 records diameter 16 with no unit. Its schema accepts a number, but semantics are unresolved. It belongs in the PF-B06 abstention segment.

PW-L007 contains an explicit dimension that conflicts with its title. Both fields have provenance. The conflict is preserved; normalization does not choose a winner.

PW-L008 is relevant by text but its permission is denied-for-assisted-use. It must never enter the candidate set. A retrieval metric computed before the permission filter can still use it for diagnostic research only if the purpose allows; product behavior cannot.

PW-L009 is authorized but stale beyond the constructed freshness policy. Chapter 6 will decide whether to exclude it or mark a degraded evidence state. It cannot appear as current fact.

PW-L010 belongs to an excluded high-consequence family. Even perfect similarity and complete attributes cannot make it in scope.

Query examples

PW-Q001 contains an exact family and explicit millimeter measurements. PW-Q002 uses descriptive language and the same evidence. PW-Q003 omits shaft diameter. PW-Q004 supplies a unitless value. PW-Q005 contains a fictional image-feature vector with ambiguous neighbors. PW-Q006 requests an excluded category.

Each query records expected behavior state independently from the expected candidate set. A system can retrieve a plausible candidate and still be required to clarify or prohibit.

Template artifact

The first vector generator assigns a large first dimension to every common-family item and a large second dimension to every long-tail item. A classifier can recover the generator template rather than meaningful visual similarity. The data-card defect is not hidden. Version 0.1.0-contaminated demonstrates it; the clean generator reduces direct template encoding and holds out templates.

The correction does not prove realistic images. It proves that the companion can detect and repair one known synthetic leakage route.

Seller-label bias

Some fixtures use phrases such as “universal fit” while structured fields remain incomplete. Seller text is preserved for lexical retrieval and threat cases but never used as the expected compatibility label. The structured trace labels it seller_stated and the behavior contract prohibits upgrading it to domain fact.

Long-tail gap

The synthetic long-tail family has too few independent templates to support a release-style comparison. The fitness report returns unsupported-segment, not a wide confidence claim. The mechanism experiment may use the cases for challenge testing but cannot report representative performance.

This record-level inspection is intentionally possible. A large future dataset still needs sampling and automated checks, but the team should retain the ability to trace representative and consequential cases back to sources and judgments.

Convert data findings into product dispositions

Data QA should change behavior scope. Use four dispositions.

Support

The data can support the stated claim for a named segment and version, subject to limitations. This is not release approval; it permits the next evidence step.

Narrow

Evidence supports only a smaller population, behavior, or mechanism. Patchwork can evaluate explicit-unit common families while excluding unitless legacy and undersampled long-tail states.

Repair and replay

A correctable process defect, such as duplicate leakage or wrong transformation, invalidates earlier results. Preserve the failed report, correct the pipeline, create a new version, and replay.

Stop

Permission, semantics, label validity, or population evidence is inadequate and no bounded repair exists. The product claim returns to the task or contract. A new model is not the next action.

Build a decision table:

Finding Data disposition Product consequence
Required unit absent narrow clarify/abstain for this state
Unauthorized source stop use exclude before retrieval output
Near-duplicate crosses release split repair and replay invalidate contaminated result
Long-tail templates insufficient narrow challenge evidence only; no release claim
Seller claim lacks domain support stop label reuse display provenance, never fit fact
Critical segments covered with valid grouped split support next experiment compare mechanisms within scope

Do not let a data team mark a defect “accepted” without the linked behavior disposition and authority. Likewise, do not ask data owners to decide product risk alone. The Applied AI Engineer connects the evidence, domain owner decides semantics, privacy/security owners decide their scope, and product/release authorities decide exposure.

Review the data card adversarially

Assign a reviewer to find statements that sound stronger than their evidence.

  • Does “sampled from catalog” hide deleted, private, or inaccessible records?
  • Does “domain reviewed” name competence, cases, disagreement, and authority?
  • Does “de-identified” explain method, residual risk, and derived artifacts?
  • Does “balanced” name the fields, weights, and population consequence?
  • Does “held out” include duplicates, templates, time, reviewer, evaluator, and provider exposure?
  • Does “synthetic” reveal generator assumptions and prohibit prevalence claims?
  • Does “approved” link a purpose, authority, scope, version, and expiry?
  • Does “ground truth” exceed what the judgment process can establish?
  • Does “representative” name the task, population, time, segment, consequence, and claim?

Rewrite every unsupported adjective as a bounded record or unknown. The card is stronger when it exposes limits than when it reads like marketing.

Finish the review by tracing one material claim backward. Start with “hybrid retrieval improves candidate coverage for the supported explicit-unit segment.” Locate the exact frozen cases, grouped split, judgment criteria, source and generator versions, permission state, transformation code, mechanism version, result, uncertainty, and limitation. Then trace forward to the only decision it supports.

If any link is absent, the claim is not yet release-quality evidence. It may remain development evidence with the gap recorded. If the trace reaches a prohibited source, contaminated partition, or mismatched criterion, invalidate the claim rather than adding a disclaimer after the result.

This backward-and-forward trace is the practical meaning of data lineage for applied AI: not an inventory diagram, but the ability to explain why one behavior claim is or is not justified.

Store the trace beside the experiment rather than reconstructing it during review. A reproducible hash without semantic lineage proves file identity, not evidence fitness; semantic lineage without immutable versions cannot prove which data actually produced the result. Both are required.

The reviewer should be able to replay that chain without relying on the original author’s memory.

PF-04 data card v0.1

Patchwork defines a local synthetic substrate.

Identity

  • Name: patchwork-catalog-synthetic
  • Version: 0.1.0
  • Seed: fixed and recorded
  • Case status: fictional/synthetic
  • Representativeness claim: none
  • Purpose: exercise candidate, permission, freshness, missingness, duplicate, leakage, and behavior-state mechanics

Population and scope

Constructed low-consequence pump-seal listing and query cases only. Excludes real buyers, sellers, marketplaces, high-consequence categories, autonomous actions, and production claims.

Record types

Listings, queries, candidate judgments, split assignments, provenance records, and data-card metadata.

Required segments

Complete/missing dimension, explicit/missing unit, common/long-tail family, fresh/stale, authorized/unauthorized, clear/ambiguous feature, unique/duplicate cluster, literal/descriptive query, single/conflicting evidence, included/excluded category.

Judgment states

Relevant, not relevant, unresolved, deterministic incompatibility, insufficient evidence, excluded. Criteria and constructed reviewer role are versioned.

Known injected defects

One near-duplicate cluster crosses the naive split; one seller label overstates fit; one image feature encodes a template artifact; one listing lacks a critical dimension; one unit is absent; one source is stale; one source is unauthorized; one long-tail family has inadequate coverage.

Prohibited uses

No claim about real prevalence, performance, fairness, user behavior, seller behavior, compatibility, safety, financial value, or release readiness. No training of a production model. No external action.

Current disposition

The naive split is rejected because duplicate and template leakage contaminate release evidence. Rebuild by canonical cluster and generation template. Long-tail and ambiguous-image gaps narrow any later claim to the constructed supported segments.

Completion exercise

Generate a seeded dataset and freeze its hash. Produce structural, semantic, and fitness reports. Identify a duplicate or correlated cluster that crosses the naive split, explain why it inflates evidence, rebuild a clean group split, and preserve both reports.

Your data card must name population, time, units, provenance, permissions, label process, reviewer competence, segments, missingness, duplicates, transformations, split roles, leakage routes, limitations, prohibited uses, retention, refresh triggers, and owner.

The exercise fails if synthetic counts become real-world claims, a data card is treated as permission, labels are called ground truth without qualification, or the clean release set influenced development.

Chapter decision

Representative data is evidence relative to a task, population, time, segment, consequence, and claim. Structure, semantics, and fitness are different validation layers. Documentation makes questions and limitations inspectable but does not prove quality or permission. Labels, missingness, duplicates, splits, and synthetic generators all carry assumptions that can leak into product behavior.

Patchwork’s PF-04 v0.1 is a fictional seeded data card and segment plan. It intentionally exposes long-tail gaps, seller-label bias, ambiguous features, missing dimensions, duplicate listings, permission/freshness states, and a contaminated split. The first split is rejected and rebuilt by canonical cluster and generator template. Unsupported compatibility claims remain narrowed.

Chapter 6 will turn the permitted, versioned records into runtime context. It will separate query transformation, candidate retrieval, ranking, permission, freshness, conflict, assembly, provenance, and empty-evidence behavior so a context failure cannot hide behind a model explanation.