← Embeddings From First Principles

Is Similarity One-Dimensional?

One cosine score tells you how a pair aligns. It does not tell you how decisive that alignment is relative to the candidates around it. Separate absolute pair score from relative-ranking and neighborhood evidence, measure what a supervised diagnostic built from that evidence actually buys on RELATE hard negatives, and keep evidence and routing policy in separate objects.

Part IV — Measuring the Representation

Two pairs, same cosine, different situations

Illustrative vignette — constructed for intuition, not a measured RELATE example.

Consider what a retrieval system typically logs when it claims a match: one number.

Pair 1:  cos = 0.78
  the query's nearest neighbor scores 0.78; the 2nd scores 0.44
  the region is sparse
  → the winner stands alone

Pair 2:  cos = 0.78
  the nearest neighbor scores 0.78; the 2nd, 3rd, 4th score 0.77, 0.76, 0.76
  the region is crowded, several near-ties within a few hundredths of the winner
  → a coin toss dressed as a match

The scalar is identical. The situations, as constructed, are not: in Pair 1 the winner has no close competition; in Pair 2 three runners-up sit within 0.01–0.02 score points of the winner and the local neighborhood is dense with candidates that could plausibly have won instead. Whatever separates a decisive match from a fragile one in this sketch was available at retrieval time. Cosine, by construction, cannot express it — not because cosine is broken, but because it was never asked the question.

One similarity score tells you how one pair aligns. It does not tell you how decisive that alignment is relative to the candidates competing around it.

Chapter 14 pushed a single scalar through an explicit calibration exercise and found a costly trade-off: at the measured near-equal-error point both class-conditional error rates were about 24%, while one particular 5%-tail abstention rule would defer 85.6% of that stress-test sample. This chapter asks the question that leaves open: was the scalar discarding evidence the representation and ranking already contained, or does this decision remain difficult even when we read more of the same space?

How much additional evidence about a retrieval decision is already sitting in the embedding space and its ranking, before a second model is ever called?

What a pairwise score actually collapses

Retrieval embeds a query q and a corpus of documents into the same space and ranks by a chosen metric — cosine, on L2-normalized vectors, is cos(q, d) = q · d. For a fixed query direction, ranking candidates by cosine is exactly reading off each candidate’s projection onto that one direction. That is precise and useful, and it is also the whole limitation in one sentence: a pairwise score collapses a candidate to a single scalar — its alignment with the query — and says nothing about where that candidate sits relative to the other candidates competing for the same rank.

But a ranking never lives at one point. It lives among competitors, and the competitors have shape: some rankings have a clear winner and a distant field; others have a crowded pack a hair’s width apart. That shape is already present in the vectors and in the full ranking the retrieval step already computed. No new semantic model is required to ask whether a result is isolated or closely contested — but, as the rest of this chapter is careful to show, some richer diagnostics require additional computation beyond reading the existing ranking, and a few require another model entirely. Those are different costs, and conflating them would undo exactly the discipline Chapter 6 installed for hubness and Chapter 14 installed for calibration.

Three layers of evidence about one retrieval decision

Sort the available evidence by what it costs to obtain, and keep two distinctions sharp throughout this chapter: conceptual diagnostics you could compute versus the specific six features Wave 1 row 1.12 actually measured, and evidence already available from the current ranking versus evidence that requires new computation.

Layer 1 — pair score. cos(q, d). What the candidate’s absolute alignment with the query looks like, in isolation.

Layer 2 — retrieval context. What the score looks like relative to its competitors and its own local neighborhood. This layer splits further by cost:

  • Derivable from the current ranking without another embedding or model call. The candidate’s rank in the full ordering; how far the candidate’s score sits below the top result (a candidate-to-top deficit, not to be confused with Chapter 11’s positive-minus-negative margin or a top-1-minus-top-2 gap — more on this below); the spread of the query’s own top-k scores. These still require lightweight derived computation, but they consume no new semantic signal.
  • Requires index-side precomputation or extra retrieval work. How crowded the candidate’s own neighborhood is among corpus items (its local density); how often the candidate turns up in other items’ neighbor lists (its in-degree, Chapter 6’s hubness).
  • Requires a perturbation protocol and additional embedding or retrieval calls — conceptual here, not measured by row 1.12. Whether the ranking survives a typo, a paraphrase, or a truncation of the query (stability); whether it responds to a meaning-changing edit such as a swapped entity or date (sensitivity). Neither of these is free: each needs a stated perturbation procedure — perturbation type, generator, seed where relevant, number of trials — recorded as part of the measurement, because a stability score without that provenance is incomplete evidence in exactly the way Chapter 11’s margin was incomplete without its negative-selection rule.

Layer 3 — an additional readout. A different decision function entirely, applied on top of or instead of the default geometry: an external model (a cross-encoder, a second encoder, a verification step) or a relation-specific learned readout fit over the same frozen embeddings. Both introduce something the base representation’s default geometry did not supply on its own — supervision, or another model’s judgment — and both are addressed later in this chapter, kept explicitly separate from Layer 2.

EvidenceLayerCostStatus here
Similarity score cos(q, d)1already computedmeasured (row 1.12: score)
Full-corpus rank2already computedmeasured (row 1.12: rank)
Candidate-to-top score deficit2already computedmeasured (row 1.12: artifact field margin)
Spread of the query’s top-k scores2already computedmeasured (row 1.12: topk_std)
Candidate’s local neighborhood density2index-sidemeasured (row 1.12: local_density)
Candidate in-degree / hubness2index-sidemeasured (row 1.12: in_degree)
Rank/ranking stability under query perturbation2perturbation protocol + callsconceptual — not measured here
Sensitivity to a meaning-changing edit2perturbation protocol + callsconceptual — not measured here
Cross-encoder / second-encoder agreement3another modelconceptual — not measured here
NLI entailment probability for (q, candidate)3another modelmeasured (row 1.12, added on top)
Verification result3another model / oracleconceptual — not measured here

Two conceptual diagnostics from earlier drafts of this idea are worth naming and then setting aside. A “participation ratio of the top-k similarities” is a real, computable quantity, distinct from anything row 1.12 measured — keep it in the conceptual toolbox, not the measured feature list. “Does the top-5 stay a subset of the top-20?” is not: under one fixed exact ranking, the top-5 is necessarily a subset of the top-20, so that question as stated carries no information. The useful version of that idea — does the neighborhood churn across a perturbed query, a different retrieval policy, or a different model — belongs to the perturbation-evidence category above, or to Chapter 16’s question about neighborhood agreement across models; it is not something this chapter measures.

The diagnostic vector

The teaching move is to stop returning a scalar and return a small structured observation instead:

match_confidence = 0.78

becomes

R(q, d) = {
  score:  0.78
  rank:   0
  candidate_to_top_deficit: 0.0     # this candidate IS the top result here
  topk_std: 0.02
  local_density: 0.81
  in_degree: 3
}

Two things about R(q, d) deserve care that the earlier vignette glossed over. First, this is an observation, not yet an interpretation: the raw fields are measurements, and turning topk_std = 0.02 into a label like LOW, or in_degree = 3 into HUB, requires a threshold — exactly the calibration discipline Chapter 14 just installed. A signal bundle that silently ships margin: 0.02 LOW without a calibration record behind that LOW is smuggling a policy decision into what is supposed to be evidence. Second, “margin” in that sketch is doing double duty across this book, and the collision is worth naming rather than hiding: Chapter 11’s margin is score(reference positive) − score(hardest selected negative); a familiar intuitive quantity is top1 − top2; and the field row 1.12 actually calls margin is neither of those — it is a candidate-to-top score deficit, computed once per candidate as candidate_score − top1_score (and pinned to 0 for the candidate that is the top result). Three different quantities, one overloaded name. The demonstration below uses the precise phrase throughout and reserves “margin” for Chapter 11’s meaning.

Reading Pair 1 and Pair 2 through this lens (still illustrative, not measured): Pair 1’s runner-up has a candidate-to-top deficit of 0.44 − 0.78 = −0.34, a large separation in magnitude, and the neighborhood is sparse. Pair 2’s next few candidates have deficits only a few hundredths below zero and the neighborhood is crowded. The same absolute top score, 0.78, therefore sits inside two very different competitive contexts. A scalar cannot express that context by itself. A small vector of relative-ranking and neighborhood observations can, and that is the claim this chapter goes on to measure.

What collapses and what does not

Not every decision needs this machinery, and the vector is a cost that has to earn its place, not a default upgrade. The governing question is not “is this application coarse or important” in the abstract; it is:

Does a scalar under a validated calibration contract (Chapter 14) achieve an acceptable error/cost profile for this decision? If yes, stop there. If no, richer diagnostics may earn their cost.

That test can cut either way, and it is worth resisting the temptation to pre-sort applications by category. “Coarse deduplication only needs a scalar” sounds safe until a false merge is expensive to undo — at which point the same “coarse” label hides a decision that needs the richer evidence after all. The right move is always to check the calibrated scalar’s measured error profile against the application’s actual cost of being wrong, not to guess from the application’s name.

Where the scalar’s calibrated operating point already meets the error budget, stop — a diagnostic vector adds cost for no measured benefit. Where it does not — retrieval feeding a downstream model for which a bad candidate can materially change the answer, record linkage on people, a safety filter — richer evidence may justify a more selective policy. The contrast is not “0.78 means accept” versus “0.78 means reject”; Chapter 14 already ruled out that shortcut. It is: given the scalar policy already calibrated for this task, does relative-ranking or neighborhood evidence improve the decisions enough to earn its additional cost?

Demonstration: does relative-ranking and neighborhood evidence beat the scalar on RELATE hard negatives?

MEASURED on RELATE v0.1, Wave 1 row 1.12 — artifact experiments/embeddings-from-first-principles/wave1/artifacts/signal-ablation.json. Model all-mpnet-base-v2; the added Layer-3 feature uses cross-encoder/nli-deberta-v3-base.

The task, exactly. For every RELATE query with at least one grade-3 positive, the implementation builds one candidate row for the first grade-3 positive (label 1) and one candidate row for every listed hard negative for that query (label 0). This produces 1,206 candidate-level examples, not 1,206 (query, positive, negative) triples — each row is a single (query, candidate, label) instance, and a classifier trained on them is learning to tell a selected reference positive apart from that query’s designated hard negatives, not choosing between a pair. The positive rate is 0.223, reflecting that most queries contribute several hard-negative rows per positive row. The hard negatives are RELATE’s designated hard-negative candidates for each query — a mix of construction methods established in Chapters 10 and 11, not a set guaranteed to score near-identically to the positive under this or any particular scorer; hardness, as Chapter 11 showed, is relative to the scorer that measures it, not a property baked uniformly into every listed candidate.

The six measured features, named precisely — artifact field versus actual meaning.

Artifact fieldWhat it actually computes
scorecos(query, candidate) — Layer 1
margincandidate_score − top1_score (0 if the candidate is the top result) — a candidate-to-top deficit, not Chapter 11’s margin and not a top-1-minus-top-2 gap
local_densitythe mean item-item cosine between the candidate and its ten nearest other corpus items — a property of the candidate’s neighborhood, not of the query’s
topk_stdthe standard deviation of the query’s own ten highest query-to-item scores — one value per query, shared across every candidate row for that query
in_degreehow many other items list the candidate among their own item-item top-10 neighbors — Chapter 6’s hubness statistic
rankthe candidate’s position in the full-corpus ranking for this query

Sorted into the layer decomposition above: score is Layer 1. rank, the candidate-to-top deficit, and topk_std are all derived from the same query-to-corpus ranking already computed — they are relative-ranking evidence, not independent new measurements. local_density and in_degree are candidate-neighborhood evidence, computed from the item-item similarity structure rather than from this particular query’s ranking at all. That distinction matters for how the result below should be read: six numbers, but not six independent channels.

Balanced accuracy, briefly. Because positives are a minority class (22.3%), the evaluation uses balanced accuracy — the average of the per-class accuracy rates — rather than ordinary accuracy, which a classifier could inflate simply by favoring the majority negative class. Balanced accuracy is a discrimination metric here, not a calibrated probability; Chapter 14 already separated those two ideas, and nothing in this experiment produces a runtime confidence score, only a classifier’s cross-validated accuracy at telling the two labels apart.

                                              mean 5-fold balanced accuracy
score only                                              0.7577
+ rank, candidate-to-top deficit, topk_std,
  local_density, in_degree  (all six features)          0.8996
+ NLI entailment probability, (query, candidate)        0.8967

geometry_lift_over_score  =  +0.1419
nli_increment              =  −0.0029

Read exactly what this establishes, and stop precisely where the evidence stops.

The six-feature set jointly supports a substantially better classifier than the score alone, under this five-fold cross-validation. That is the whole, correctly bounded claim. It is not evidence that all six features are independently useful — there is no one-feature-at-a-time ablation, no correlation analysis, and no permutation-importance result in this artifact, and two of the six (rank and the candidate-to-top deficit) are themselves transformations of the same score ordering rather than separate sources of evidence. It is also not evidence about where the gain came from: the artifact records no comparison of which specific rows moved from wrong to right, so the claim that the lift falls “entirely on the cases score alone got wrong” is not something this experiment measured — a classifier’s balanced accuracy can rise through some combination of correcting old errors and introducing new ones on different rows, and distinguishing those requires a prediction-overlap analysis this chapter does not have.

The cross-validation is not grouped by query. cross_val_score(..., cv=5) was called without supplying query identifiers as groups, so rows from the same query are not guaranteed to stay together across folds. The correct description of 0.8996 is mean five-fold candidate-level cross-validated balanced accuracy under this implementation — not “held-out query” performance and not evidence of generalization to genuinely unseen queries. A query-grouped rerun, keeping every candidate row from one query in a single fold, is the natural next experiment to test whether the lift survives that stricter split; it has not been run, and this chapter does not guess its result.

Adding the NLI feature did not improve the reported mean. 0.8967 against 0.8996 is a measured difference of −0.0029. No fold-level scores, standard deviation, or confidence interval were persisted for either run, so this difference cannot be called “noise,” “within noise,” or “statistically insignificant” — none of those phrases are supported without an uncertainty estimate this artifact does not contain. The correct statement is exactly this narrow: adding this NLI feature did not improve the reported mean balanced accuracy in this experiment. Chapter 10 already gave a specific, independent reason to be cautious of this particular NLI model’s (query, candidate) entailment-probability readout — its orientation and objective were not validated as a support score across RELATE’s mixed query styles — but row 1.12 does not isolate why the increment came out negative here, and this chapter does not adopt that Chapter 10 caution as a causal explanation for this particular result.

Limitations worth holding in view together, because several of them compound: the positive class is the first grade-3 positive only, not every acceptable answer; the hard negatives are RELATE’s designated set, mixed in construction method; the unit of analysis is a candidate instance, not a query-level decision; the cross-validation is not query-grouped; no fold-level uncertainty exists for either the geometric lift or the NLI increment; and no individual feature’s contribution was isolated. None of that empties the result — a +0.1419 mean lift from the six jointly supplied geometric/ranking features on this candidate-level classification task is a substantial measured difference. It is a narrower finding than “the geometry diagnoses its own uncertainty entirely on the score’s own mistakes,” and the narrower version is the one this chapter defends.

A precise way to hold the whole result: the fixed embedding space plus the query ranking expose relative-ranking and candidate-neighborhood structure that a supervised classifier can use to separate this query’s reference positive from its hard negatives substantially better than the pairwise score alone allows. The relative-ranking features come from the query-to-corpus ordering; local_density and in_degree come from the item-item geometry. The geometry did not decide this on its own — a logistic regression, trained on labels, learned how to weigh six derived features. That is a meaningfully weaker and more accurate claim than saying the geometry “diagnoses its own uncertainty,” and it is the one the evidence actually supports.

A stronger question: what if we change the readout entirely?

Everything so far still reads the embedding’s default geometry: cosine reads it with one scalar; the six-feature classifier reads more of the same geometry — its ranking and its neighborhood structure — but every one of those six numbers is still downstream of the encoder’s own objective and its default distance function. There is a further move available: keep the embedding frozen, but fit a small readout aimed at one specific relation, and measure distance in that readout’s coordinates instead of in cosine space.

cosine              one scalar over the default geometry
   |
diagnostic vector   read more of the same geometry — relative ranking, neighborhood
                     (this chapter, row 1.12: +0.1419 mean balanced accuracy)
   |
learned readout     fit a projection for one chosen relation, then measure
                     distance in that projection's coordinates

The RELATE project’s RelationProjection mechanism is exactly this third move. The result below is preserved as historical external evidence in RELATE’s evidence ledger; it is not a RELATE-corpus Wave artifact, and the original frozen CodeBERT/AST assets are not present in the current replay environment for digit-level recomputation. RELATE preserves the historical figures and separately maintains replay/mirror machinery, so the provenance boundary should travel with the numbers rather than disappearing behind them.

EXTERNAL HISTORICAL RESULT — not a RELATE-corpus Wave artifact and not freshly recomputed in the current environment. Frozen CodeBERT embeddings of roughly 20,000 Python functions (CodeSearchNet); the relation is code structure — cyclomatic complexity, maximum control-flow depth, and distinct call sites, read from the abstract syntax tree. Preserved by the RELATE project’s evidence ledger — repository, interactive demo.

The task: pairwise structural ordering — given a query function and two candidates, name the candidate structurally closer to the query. Same encoder, same frozen 768-dimensional vectors, 4,000 test queries with 128 hard-negative comparisons each. Chance is 0.500.

method                                    pairwise ordering accuracy
token-length difference                   0.499
raw cosine distance                       0.532
raw Euclidean distance                    0.533
ridge readout -> 3 structure coordinates  0.733
ground-truth-coordinate oracle            1.000

A ridge projection from the frozen embedding into three predicted structure coordinates, ranked by Chebyshev distance over the robust-scaled predictions, ordered the structurally closer candidate 73.3% of the time — an absolute gain of about 19.95 percentage points over the stronger raw-distance baseline (Euclidean, 0.533). The predicted coordinates are far from perfect; the gap to the 1.000 row — an oracle ceiling built from the same ground-truth AST coordinates that define the target relation, not a universal theoretical bound any representation should be expected to approach — stays large. The token-length baseline sitting at chance argues against the simplest possible confound: that the readout is merely tracking function length. It does not rule out every possible correlate of program size or structure a more elaborate baseline might expose; it rules out the trivial one.

The caveat that makes this honest. The readout was trained on AST-derived structural coordinates from the training split. Cosine and Euclidean distance received no information about the relation at all. This is not an apples-to-apples contest between two unsupervised metrics — one side was given supervision the other was denied. The narrower claim it supports:

A frozen embedding can hold recoverable information about a chosen relation that its default geometry exposes poorly, and a small supervised projection can recover enough of it to materially improve hard-negative ordering — for this relation, this representation, and this readout.

That is the same shape as this chapter’s own row 1.12 result, one layer deeper: the diagnostic vector showed that a scalar discards decision-relevant ranking structure; the learned readout shows that the default distance function itself is not the representation’s only useful readout. It echoes Chapter 7’s point that a vector can carry more than its default readout exposes, and it anticipates a move Chapter 21 develops much further — a supervised bridge that can beat source-native cosine on hard negatives — without pre-teaching that chapter here.

Where this stops. A learned readout does not always help, and this book does not let that caveat go unmeasured. On PAWS, a paraphrase benchmark from the same research programme, an alternative embedding-pair readout scored below cosine (0.6205 → 0.5945); on BigCloneBench it scored well above it (0.7255 → 0.856). “There is always more structure than cosine shows” is too strong a generalization from that spread. Whether a relation is recoverable from a given representation, and whether a given readout recovers it, are empirical questions — per relation, per representation, per readout — exactly the discipline this book applies to every other geometric claim. Recoverability is not a property of embeddings in general; it is something to measure for the case in front of you.

The demo lets you compare cosine distance, readout distance, and true AST distance on your own Python functions. The live examples illustrate the mechanism; they add no benchmark evidence beyond what is reported above, and a single good or bad example neither confirms nor overturns the frozen result.

What this chapter establishes and what it does not

Establishes: an absolute pairwise score discards information about a candidate’s rank, its deficit relative to the top result, its own local neighborhood density, and its in-degree — all already latent in a ranking the retrieval step already computed, distinct from the extra cost of perturbation- or external-model-derived evidence; on RELATE v0.1, a classifier built from six such features achieves mean five-fold balanced accuracy 0.8996 against 0.7577 for the score alone, a joint lift of +0.1419, on 1,206 candidate-level examples (positive rate 0.223) built from the first grade-3 positive per query against that query’s designated hard negatives; adding an NLI entailment feature changes the reported mean to 0.8967, a −0.0029 difference; and the RELATE project’s RelationProjection mechanism, evaluated externally on frozen CodeBERT embeddings, shows a relation-specific supervised readout can substantially outperform raw cosine and Euclidean distance for one chosen relation, without proving learned readouts are generally superior.

Does not establish: that any one of the six row-1.12 features is independently useful (no per-feature ablation exists); that the measured lift comes specifically from cases the score-only classifier got wrong (no prediction-overlap analysis exists); that the result generalizes to genuinely unseen queries (the cross-validation was not query-grouped); any statistical significance for either the +0.1419 lift or the −0.0029 NLI change (no fold-level uncertainty was stored); why the NLI feature failed to help (the experiment does not isolate a cause); that a diagnostic vector is always worth its cost, or that a scalar is always insufficient; that this experiment measured perturbation stability, semantic sensitivity, participation ratio, or query-side density — none of these were computed here; or that a learned readout beats cosine in general (the PAWS counterexample stands against that).

Lab 15: reproduce the measured lift, and its exact boundaries

MEASURED — artifact experiments/embeddings-from-first-principles/wave1/artifacts/signal-ablation.json (mpnet-base, plus cross-encoder/nli-deberta-v3-base for the third feature set). REPRODUCIBLE — python run_wave1.py 1.12.

Question. How much of a hard-negative decision is already recoverable from evidence the ranking already contains, before any second model is called — and exactly how far does that evidence go?

Step 1 — define the candidate-level task. For each qualifying RELATE v0.1 query: selected positive = the first grade-3 positive; negative examples = every listed hard negative for that query. One row per (query, candidate). Measured: n_examples = 1206, positive_rate = 0.223.

Step 2 — compute the score-only feature and classify. score = cos(query, candidate). Balanced logistic regression, five-fold cross-validation. Measured: 0.7577.

Step 3 — compute the exact six-feature geometric set. score; the artifact’s margin field, which is candidate_score − top1_score and 0 for the top-ranked candidate — call it the candidate-to-top deficit in your own notes; local_density, the mean item-item cosine between the candidate and its ten nearest other corpus items; topk_std, the standard deviation of the query’s own top-10 scores (identical across every candidate row of that query); in_degree, the candidate’s item-item hubness count; rank, the candidate’s position in the query’s full ranking. Measured combined result: 0.8996 (+0.1419 over score-only).

Step 4 — interpret only the joint result. Do not attribute the gain to any one feature; no individual ablation was run.

Step 5 — add the external NLI feature. Entailment probability for (query, candidate). Measured: 0.8967 (−0.0029 from the geometric set). State plainly that it did not improve the reported mean; do not add a cause.

Step 6 — audit the cross-validation. cross_val_score(..., cv=5) was called without query groups. Record this as a scope limitation on every claim drawn from 0.8996 or 0.8967.

Step 7 (PROPOSED — no artifact backs this) — query-grouped replication. Re-run the same feature sets with every candidate row from one query constrained to a single fold. Measure whether the +0.1419 lift survives.

Step 8 (PROPOSED — no artifact backs this) — feature ablation. One feature at a time; nested subsets; leave-one-feature-out; optionally permutation importance. Measure which of the six features carry incremental information once the others are already present, since rank and the candidate-to-top deficit are derived from the same ordering and may not add independently.

Step 9 (PROPOSED — no artifact backs this) — inspect changed predictions. Compare the score-only and six-feature classifiers row by row: which cases flipped from wrong to right, which flipped the other way, and whether either group concentrates by relation type, query style, or rank/density strata. This is the analysis that would be required before claiming the lift falls “entirely” on the score’s own failures — it has not been run.

Try it yourself

Build the same candidate-level task on your own hard cases: one row per (query, candidate), score plus rank, candidate-to-top deficit, top-k spread, local density, and in-degree. Fit the score-only and six-feature classifiers, using query-grouped cross-validation from the start rather than adding it as an afterthought. Inspect ten cases where the richer classifier is right and the scalar is wrong; check whether they look like near-ties, hubs, or fragile rankings, or something else in your data. Decide, from the measured lift and its cost, whether your application needs the full vector or the scalar plus a calibrated threshold is enough.

Companion component: the signal bundle, and a separate routing policy

Evidence and action are different objects — Chapter 6 established the general rule, Chapter 12 carried it into retrieval policy, Chapter 14 carried it into calibration, and this chapter’s bundle should not regress it by putting a verdict field inside the evidence record.

signal_bundle:
  provenance:
    space_record_ref:
    retrieval_policy_ref:
    feature_definition_version:      # a field named "margin" must not silently change meaning
    query_id:
    candidate_id:

  pair_score:
    score:                           # Layer 1

  ranking:                            # Layer 2, already available from the current ranking
    rank:
    candidate_to_top_deficit:
    topk_score_std:

  candidate_neighborhood:             # Layer 2, index-side
    local_density:
    in_degree:

  perturbation_probes:                # Layer 2, conceptual unless run
    stability:
    sensitivity:
    perturbation_spec:
    status:            <not_run | measured>

  external:                           # Layer 3
    nli_entailment:
    second_encoder_agreement:
    verification:
    status:            <not_run | measured>

routing_policy:
  id / version:
  consumes:            signal_bundle
  action:               <accept | rerank | requery | verify | abstain>

The split is deliberate and mirrors the rest of the book: the bundle records what was measured; the routing policy decides what to do about it. feature_definition_version exists because this chapter just demonstrated a real collision — a field called margin means something different here than it does in Chapter 11, and a runtime that lets that ambiguity travel silently between two components is one accidental refactor away from comparing incompatible numbers under the same name. status fields on the perturbation and external blocks exist so that a signal bundle never implies a perturbation probe or an external call happened when it did not; an unrun probe is a not_run status, not a missing number quietly treated as zero.

Embedding Observatory behavior

The Observatory should not compute every candidate diagnostic on every retrieval, and it should not describe any of them as free. A defensible cost ladder, cheapest first:

existing ranking evidence (score, rank, candidate-to-top deficit, top-k spread)
   → index/corpus diagnostics (in-degree, local density) — precomputed or looked up
      → perturbation probes (stability, sensitivity) — sampled traffic or evaluation-time only
         → external semantic model (NLI, second encoder, verifier) — paid escalation
            → human or high-cost verification

Layer 1 and the ranking-derived part of Layer 2 require no additional embedding or external-model call, only lightweight computation over scores the retrieval already produced. Candidate-neighborhood diagnostics typically need item-item precomputation or lookup. Perturbation probes need extra embedding and retrieval calls and should be run on sampled traffic, at evaluation time, or when a routing policy specifically requests them — not on every query by default. External semantic calls are a further cost tier and should be invoked only when a validated routing policy shows that the cheaper evidence is insufficient for the decision.

The Observatory should also version its feature definitions and record which version produced any stored bundle. This is not bureaucracy for its own sake: this chapter’s own row 1.12 margin field, if stored without that version tag next to a Chapter 11 margin computed from the same corpus, would be a silent, undetectable semantic collision the next reader has no way to catch.

Failure modes

  • Calling every relative-score quantity a “margin.” Tempting because they all involve subtracting two scores. Check: Chapter 11’s positive-minus-negative margin, a top-1-minus-top-2 gap, and row 1.12’s candidate-to-top deficit are three different quantities; name the one you mean.
  • Treating “geometric” as “free.” Tempting because no second model is called. Check: candidate-neighborhood diagnostics need index-side work, and perturbation probes need new embedding calls; only the already-computed ranking evidence is genuinely free.
  • Converting a detector into a verdict. Tempting because a high in-degree or a flat response to an edit looks like a clear signal. Check: hubness, low stability, and insensitivity are evidence to investigate, not automatic proof of a bad match — Chapter 6’s rule again.
  • Treating six features as six independent signals. Tempting because the table lists six rows. Check: rank and the candidate-to-top deficit are both derived from the same score ordering; no ablation in this chapter shows independent contribution.
  • Reporting candidate-level cross-validation as query-level generalization. Tempting because the number sounds like held-out performance. Check: no query groups were supplied to the cross-validation here; replicate with grouped folds before claiming unseen-query evidence.
  • Calling −0.0029 “noise.” Tempting because it is a small number. Check: no fold-level uncertainty was stored for either feature set; “noise” is a claim this artifact cannot support.
  • Explaining the NLI decrement causally. Tempting because Chapter 10 already gave a plausible mechanism for this scorer’s weaknesses. Check: that mechanism is a reason for caution about the scorer, not a demonstrated cause of this specific −0.0029.
  • Putting the routing verdict inside the evidence bundle. Tempting because it is convenient to ship one object. Check: evidence and policy are different objects with different owners and different validation requirements.
  • Adding expensive diagnostics where a calibrated scalar already meets the application’s error budget. Tempting because more signal always sounds safer. Check: complexity has to earn its cost against a measured requirement, not against a vague sense of importance.

What this chapter established

  • An absolute pairwise score discards context already available in the same retrieval system: relative-ranking evidence (rank, candidate-to-top deficit, top-k spread) from the query-to-corpus ordering, and candidate-neighborhood evidence (local density, in-degree) from the item-item geometry. Perturbation probes and external-model evidence are separate, more expensive layers.
  • Row 1.12’s artifact field margin is candidate_score − top1_score — not Chapter 11’s margin, and not a top-1-minus-top-2 gap. Naming collisions like this are exactly why feature definitions get versioned.
  • Measured: six jointly-supplied features lift balanced accuracy from 0.7577 to 0.8996; adding an NLI feature moves it to 0.8967. The lift is joint — not evidence that each feature is independently useful, that the gain falls on the score’s own failures, or that it generalizes to unseen queries, since the cross-validation was not query-grouped.
  • A supervised classifier over ranking and neighborhood features is a learned readout of the geometry, not the geometry autonomously diagnosing its own uncertainty.
  • The signal bundle records evidence tagged by layer, cost, and status; a separate routing policy consumes it and chooses an action. Evidence and policy stay different objects, all the way down.

Next

Part IV progressively separated the representation from the instruments used to judge and act on it: Chapter 13 fixed an evaluation contract so representations could be compared; Chapter 14 bound an operating point to a calibration contract; and this chapter held one space fixed while reading more of its ranking and neighborhood structure, then showed an external historical example of a supervised relation-specific readout over frozen vectors. Part V now removes the fixed-space assumption entirely: if the encoder itself changes, what—if anything—remains comparable? Part IV asked how to measure and read an embedding space. Part V asks what survives when the space itself changes.