NEWProduction Web Themes & Turnkey ArchitecturesGet Lifetime Pass ($199) →
KNKomal Nakrani
Get All Access
ThemesDocsAll-Access PassGet All Access ($199)
All books
LLM Engineering
LLM Behavior Engineering
Contracts, Context, Retrieval, and Evaluation
Komal Nakrani
First edition/LLM Engineering · Volume 1

LLM Behavior Engineering

Contracts, Context, Retrieval, and Evaluation

A professional guide to making language-model behavior explicit, grounded, measurable, and change-ready across managed and open-weight systems.

Edition
First edition
Version
1.0.0
Published
Table of contents
01The Model Is Not the BehaviorSeparate a model response from configured system behavior, outcome evidence, and accountable human authority.65 min02Write the Language-Task ContractTurn an ambiguous language feature into measurable states, consequences, non-goals, escalation, and named decision authority.75 min03Reason About Tokens, Attention, and GenerationUse bounded mechanism knowledge to predict token, template, context, and decoding failure surfaces without anthropomorphism.80 min04Select a Model and Access PostureChoose a reversible model and access path from task evidence, hard constraints, responsibility, and lifecycle risk.80 min05Establish a Reproducible BaselineBind every behavior-facing input and version into a replayable baseline with raw evidence, trials, and explicit unknowns.85 min06Engineer Instructions and MessagesBuild a versioned message interface that keeps trusted control separate from untrusted data and tests instruction changes by ablation.90 min07Make Outputs Typed and BoundedTurn generated language into a versioned proposal interface with layered validation, provenance checks, and explicit terminal states.90 min08Budget Context DeliberatelyAllocate finite context among control, authorized evidence, history, and output while making every omission and failure state visible.90 min09Design the Retrieval QuestionDecide whether external evidence is needed and specify authority, permission, freshness, evidence units, and no-evidence behavior before choosing retrieval technology.90 min10Build the Retrieval and Reranking PathTurn an authorized evidence corpus into a measurable retrieval funnel whose ingestion, chunking, filtering, ranking, and reranking decisions remain separately traceable.95 min11Assemble Context With ProvenanceConvert selected retrieval candidates into an authorized, attributable evidence bundle that preserves source state, trust, qualifying spans, and bounded behavior.95 min12Evaluate Retrieval and Generation TogetherDiagnose corpus, retrieval, assembly, generation, citation, and evaluator behavior independently and jointly without turning one score into causal evidence.100 min13Build Representative Evaluation CasesTurn intended use into a documented, segmented, leakage-aware evaluation asset whose coverage and gaps remain explicit.100 min14Name Errors and Judgment MethodsTurn raw failures into a consequence-aware taxonomy and calibrate deterministic, model, human, domain, and authority judgments without hiding disagreement.105 min15Run Experiments That Isolate ChangeTurn a proposed system change into a preregistered, paired, repeatable comparison with visible segments, uncertainty, confounds, and explicit disposition.105 min16Release, Observe, and Change the ModelTurn the complete behavior dossier into an authority-bound release, observation, rollback, migration, and adaptation-referral system.115 min