← Hallucination From First Principles

Appendix: The Evidence Ledger

Every empirical claim in the book in one place: what was measured, on which run, with what provenance, and what the number does not establish. The book applies its own discipline to itself.

The book argues that a claim should be governed by its provenance, its evidential support, and an explicit statement of what it does not establish. This appendix applies that discipline to the book itself.

Every original empirical number produced for this book is listed here with three things: where it was measured, what its provenance status is, and what it does not establish.

Two facts are visible immediately from the ledger:

  1. The Empirical Core: Every original experimental number produced for this book comes from Chapters 5 through 7 and traces to two sources: one 10,000-row HaluEval summarization run, and one three-domain hard-negative table. Chapter 8 interprets those results; it does not add new original model-performance measurements.
  2. The Architectural Derivation: Chapters 9 through 15 contain no original empirical model-performance measurements. The consistency and sensitivity axes, the abstention router, the policy engine, the recovery loop, and the memory architecture are derived from the earlier results and from first principles. The runnable demonstrations in those chapters are deterministic tests of architectural invariants under a perfect structured oracle. They establish that the types hold. They establish nothing about the accuracy of a real verifier or model.

A reader who finishes Chapter 15 should not believe the architecture has been validated. It has been reasoned through. The experiments that would validate it are listed in Table B.

    flowchart TD
    A[Raw Data: HaluEval 10k Run] --> B[Feature Extraction: Geometry, Entailment, Structure]
    B --> C[Measured Results: Chapters 5-7]
    C --> D[Interpretation & Ceilings: Chapter 8]
    D --> E[Derived Architecture: Chapters 9-15]
    E --> F[Owed Experiments: Table B]
    F -.->|Future Validation| E
  

The appendix separates measured results from derived architecture: Chapters 9 through 15 consume earlier evidence but still owe direct validation experiments.

A. What the book measured

All Hallucination Energy results trace to the Certum implementation and the config-preserved summarization run 20260216_225831.

A.1 — HaluEval summarization (Chapter 5 §16, §18; Chapter 6)

Run configuration (recorded in the run’s config.json): HaluEval summarization, $N = 10,000$ rows; embedding model sentence-transformers/all-MiniLM-L6-v2; NLI model MoritzLaurer/deberta-v3-base-mnli-fever-anli; geometry rank $r = 32$, $top\_k = 1000$; run seed 1337; sentence-decomposed pipeline; 70/30 stratified row split plus 5-fold stratified cross-validation; modeling seed 42. Re-running the modeling stage reproduces every value below.

ClaimWhereProvenanceWhat it does NOT establish
Geometry feature bundle: AUC $0.7120$, 95% CI $[0.6929, 0.7306]$, 5-fold CV $0.7148$Ch5 §16Reproduced from run 20260216_225831that the single Hallucination Energy scalar carries this signal — this is a 13-feature bundle (similarities, coverage, energy aggregates); source-, entity-, or embedder-independent generalization
Entailment feature bundle: AUC $0.6126$Ch5 §16ReproducedNLI’s standalone value on any other task
Geometry + entailment + structure, logistic: AUC $0.7284$, 95% CI $[0.7108, 0.7455]$, average precision $0.7389$Ch5 §16Reproducedthat the combination transfers; the incremental contribution of $H$ specifically
Full-feature XGBoost, held-out row split: AUC $0.7338$Ch5 §16Reproducednon-row-level generalization
Full-feature XGBoost, 5-fold stratified CV: AUC $0.7515 \pm 0.0018$Ch5 §16Reproducedsource-document-grouped cross-validation (this is row-level)
Feature ablation: full $0.7284$; without energy_gap $0.7284$; without high_energy_count $0.7271$; without both $0.7271$Ch5 §18Reproducedthat the removed features are useless, or that $H$ is indispensable — related geometric substitutes remained in the bundle
Rank sweep (geometry bundle held-out AUC): $r = 1 \to 0.7211$, $r = 2 \to 0.7147$, $r = 4 \to 0.7082$, $r = 8 \to 0.7115$, $r = 16 \to 0.7113$, $r = 32 \to 0.7117$, $r = 64 \to 0.7117$. The $r = 32$ row reproduces the row above ($0.7117$ held-out; $0.7120$ bootstrap mean).Ch5 §5New finalization work, same frozen configurationsingle-scalar-vs-rank behaviour; false-acceptance-rate-vs-rank; cross-embedder behaviour. The bundle is nearly flat from $r = 4$ to $r = 64$ — rank is not a lever for this workload

A.2 — Hard-negative separation, three domains (Chapter 5 §20; Chapter 6 §5; Chapter 7 §12; Chapter 8 §15)

Negative construction: hard_mined_v2 — for each claim, shortlist the 16 evidence sets with the highest centroid similarity, exclude the claim’s own evidence and any set sharing its source page, and select the most similar remaining mismatch. Selection is by similarity, not energy: this is a semantic-neighbour hard negative (difficulty Level 3), not a metric-aware adversary. Recorded in adversarial_consolidated.csv (rerank_by_energy: false). Seed 1337.

DomainThreshold $\tau$Supported acceptanceFalse acceptanceSupported mean energyHard-neg mean energyMean gapStandardized separation
Wikipedia / FEVEROUS-style0.14680.0670.0060.37140.6950+0.3236+1.92
PubMed0.43820.6440.0140.37000.6944+0.3244+2.11
CaseHOLD0.38620.0290.0120.66800.6075−0.0605−0.40
    flowchart LR
    subgraph Factual Domains
        W[Wikipedia / FEVEROUS] -->|Separation +1.92| P[PubMed]
        P -->|Separation +2.11| S[Success: Thresholding Works]
    end
    subgraph Relational Domain
        C[CaseHOLD] -->|Separation -0.40| F[Failure: Energy Inverts]
    end
  

What these do not establish:

  • Wikipedia / PubMed: an operating point at any other false-acceptance budget; behaviour under a metric-aware adversary or a larger search budget $K$; that the strict-threshold supported acceptance (6.7% for Wikipedia) is usable in production.
  • CaseHOLD: that legal language is intractable for embeddings, or that role binding is the only missing variable. Chapter 8 §15 states the bounded reading: under this representation, evidence construction, rank, and hard-negative regime, containment energy does not preserve the ordering required to separate supported from structurally similar legal alternatives. The detector inverts on this task; a threshold cannot fix a sign flip.
  • Note: The standardized separation is a distributional quantity (a Cohen-style ratio). It is not an operating point.

A.3 — The negative result (Chapter 5 §22)

The search for a second scalar that would recover the missing separation under hard adversarial overlap — participation ratio, projection ratios, leave-one-out sensitivity, similarity margins, adaptive monotone combinations, gap-width tuning — did not reliably succeed. This is reported as a negative result, not omitted.


B. What the book owes

These experiments are named in the chapters as not yet run. They are not fabricated in prose.

ExperimentWhereRunnable now?What it would settle
Single-Hallucination-Energy-scalar AUC (all ranks)Ch5 §16, §18, §27; Ch6 §22Yes — the feature pipeline existshow much discrimination is the $H$ scalar itself vs the geometry bundle
Max / mean cosine and centroid-distance baselinesCh5 §27; Ch6 §22Yes — features already extractedwhether SVD containment beats simple proximity
Source-document-grouped, entity-disjoint, semantic-cluster splitsCh5 §17, §27; Ch6 §11, §21Needs a source-document id in the rowswhether the AUC survives isolation by source / entity / semantic neighbourhood
Embedding-model sweepCh5 §27; Ch6 §22Partiallywhether the signal transfers across sentence encoders
Centered vs uncentered geometryCh5 §11, §27Yes — one-line changewhether the detector uses evidence-local structure or global embedding anisotropy
Energy-aware hard-negative mining (hardest_energy_mined)Ch5 §19, §27; Ch7 §5, §7Needs a small Certum change (the calibration path hardcodes the similarity-only mode)whether the metric-aware adversary degrades separation further than the semantic-neighbour negatives
Adversarial degradation curve ($K = 1, 4, 16, 64, 256$)Ch6 §14; Ch7 §7Yeshow fast robustness degrades as the adversary is allowed to search harder
Projection coordinates $z$ vs scalar $H$ on structural adversariesCh8 §11, §20.AYes, with an oracle-controlled corpus (the per-example basis makes it subtler than a raw feature stack)whether the scalar reduction $\|z\|^2$ discarded structural signal that survived into the coordinates
Structured mutation response; encoder/pooling transfer; structural-extractor evaluation; evidence-size recombination stressCh8 §20.B–EMixedthe mechanism and generality of the structural ceiling
Perturbation-fidelity experiment (paired invariance / responsiveness rates, direction-correct rate, effect-vs-noise)Ch9 §20Needs the prompt families and repeated samplingwhether the consistency and sensitivity axes carry signal, and how independent they are
Contamination metrics — taint-escape rate, exposure before containment, descendant count per contaminated root, propagation depth, post-revocation exposure, false invalidation rateCh14 §17Needs a memory testbedhow much contaminated state escapes, and how fast

C. The runnable demonstrations (Chapters 12–15)

These are deterministic scripts, standard library only. They exercise the architecture’s type discipline. They are not model-quality experiments.

ArtifactWhereWhat it demonstratesWhat it does NOT establish
distrustful_controller.pyCh15The assertion, action, and persistence gates give different answers for the same content: cand_v2 (one authoritative filing, no recorded period) is authorized as a response but quarantined from durable factual memory; cand_v3 (a second independent source family plus a validity window) is admitted. An external send stays blocked until a scoped human-review artifact exists.that a real support / provenance / epistemic-adequacy check would classify cand_v2 this way — measure() here is a perfect oracle
recovery_demo.pyCh13 §18Omission terminates at PERMIT; a hedge that changes wording without changing evidence state repeats and terminates at ABSTAIN via cycle detection; an undeclared repair mode is rejected before any repair runsthat a real verifier detects the failed claim, or that a real repair model produces a sound edit
policy_engine.pyCh12 §11, §15A policy decision is a typed record (commitment, next action, escalation, obligations, reason codes), not a flat enum; replaying one fixed record under two policy versions changes the route without regenerating the candidateanything about the reliability records the engine consumes — they are synthetic
    flowchart TD
    V[Diagnostic Vector: H, C, S, E, A] --> C1[Consumed as Settled Object from Ch 12]
    C1 --> C2[Axes & Independence: NOT VALIDATED in Book]
    C2 --> C3[Requires: Perturbation-Fidelity Experiment]
  

The diagnostic reliability vector $(H, C, S, E, A, \dots)$ is consumed as a settled object from Chapter 12 onward. Its axes, their combination rule, and their statistical independence are not validated anywhere in the book (Chapter 9 §18 says so explicitly). The perturbation experiment in Table B is what would validate it.


D. External results the argument leans on

The book cites primary sources for its research claims. The load-bearing ones:

Claim used in the bookSourceWhat it establishes
Evaluations reward guessing over abstention; under binary grading, an overconfident guess is optimal and “I don’t know” is maximally penalizedKalai, Nachum, Vempala, Zhang, Why Language Models Hallucinate (OpenAI, 2025; arXiv:2509.04664)the abstention argument in Chapter 1 §5 and Chapter 11
Hallucination does not fall monotonically with capability — on OpenAI’s PersonQA evaluation the reasoning models o3 and o4-mini hallucinate at 0.33 and 0.48 against o1’s 0.16OpenAI, o3 and o4-mini System Card (April 2025)Chapter 1 §8
Semantic entropy: cluster generations into meaning classes by bidirectional entailment, then compute entropy over clusters; it targets confabulations, not systematic errorKuhn, Gal, Farquhar (ICLR 2023); Farquhar, Kossen, Kuhn, Gal, Nature 2024Chapter 4 §12
FEVEROUS: 87,026 claims with Wikipedia sentence and table-cell evidence, labelled SUPPORTS / REFUTES / NOT ENOUGH INFO — labels relative to the annotated evidence, not world truthAly et al., NeurIPS 2021Chapter 3 §15, Chapter 5
Conformal risk control: for a non-increasing bounded loss, $\hat{\lambda} = \inf\left\{\lambda : \frac{n}{n+1}\hat{R}(\lambda) + \frac{B}{n+1} \leq \alpha\right\}$ gives $\mathbb{E}[L(\hat{\lambda})] \leq \alpha$ for a new exchangeable pointAngelopoulos, Bates, Fisch, Lei, Schuster, Conformal Risk Control (2022)Chapter 6 §9
Conformal factuality: a back-off procedure removes or generalizes claims until a high-probability factuality bound holdsMohri, Hashimoto (ICML 2024)Chapter 6 §9, Chapter 13 §23
Reasoning fine-tuning degrades abstention by ~24% on average; scale gives little improvementKirichenko, Ibrahim, Chaudhuri, Bell, AbstentionBench (2025)Chapter 11 §19
An index that penalizes hallucination and rewards abstention does not rank the highest-accuracy models first — they lose points by guessingAA-Omniscience (Artificial Analysis, 2025; arXiv:2511.13029)Chapter 11 §10
Bitemporal modelling: valid time (when a fact was true) vs transaction / system time (when it was recorded) is the standard modelSnodgrass, Developing Time-Oriented Database Applications in SQL; SQL:2011Chapter 14 §12
Label noise measurably limits detector scores on a RAG hallucination-detection benchmarkChen et al., ACL 2026Chapter 6 §16

Full citations are in each chapter’s Research roots.


E. The honest summary

  • Measured: that a bundle of geometry-derived features separates supported from unsupported summary sentences on one HaluEval run at roughly $0.71$ AUC, rising to about $0.73\text{–}0.75$ with entailment and structural features; that this signal survives semantic-neighbour hard negatives in two factual domains; that it inverts on a relation-heavy legal task; that retained rank barely matters for the bundle on this workload.
  • Derived, not measured: everything about consistency, sensitivity, abstention routing, policy, recovery, and memory governance — Chapters 9 through 15.
  • Owed: the isolated-scalar baseline, the simple-proximity baselines, dependency-correct splits, cross-embedder transfer, the energy-aware adversary, the coordinate-vs-scalar test, the perturbation-fidelity experiment, and the contamination metrics.

The destination of the book is a method, not a validated system. This appendix is the book keeping its own score.