← Embeddings From First Principles

The Nearest Neighbor Can Be Wrong

A vector score can prefer a candidate that is on-topic, lexically similar, geometrically close, and wrong in polarity, role, time, or claim. Name the ways this happens, then use RELATE typed hard negatives to measure whether cosine — or a second-stage entailment score — separates a known-good candidate from the near miss.

Part III — Retrieval Is an Experiment

The result that is closest and wrong

Illustrative example — not a RELATE measurement.

Query:

Is Dublin the capital of Ireland?

Imagine the passage that ranks first is this one:

Dublin is not, and has never been, the capital of Ireland — that
distinction belongs to the older seat of government at Tara.

Fluent, on-topic, confidently phrased, and among the geometrically closest things in the corpus — and false. A model handed this passage as context may repeat its claim. The retrieval system did its job: it found a near vector. “Near” was not “correct.”

Chapter 9 removed one suspect for that measured workload. Its HNSW experiment reproduced the exact top 10 with at least 0.9937 mean set overlap, so approximation error explains very little of that particular run. That sharpens the question this chapter asks: what if the vector scoring rule itself prefers the wrong candidate? Even a perfectly faithful index can reproduce an exact ranking and still hand back a candidate that is topically appropriate, lexically similar, geometrically close — and wrong in polarity, role, time, scope, or claim.

The index can be right about the geometry while the geometry is insufficient for the decision you actually need.

Why would the geometry do this? General-purpose embedding similarity often rewards broad semantic relatedness — call it aboutness — even when assertion-level details such as polarity, temporal scope, quantity, or entity roles differ. Negating a sentence, swapping its arguments, changing a year, or altering one claim while preserving the topic can leave much of the text and subject matter unchanged. That makes these natural stress tests for a scalar similarity readout: will the changed assertion move far enough to rank below a known-good candidate?

One caution the book has raised before (Chapters 3 and 4): a cosine that fails to separate a distinction does not tell you the model never encoded that distinction. The information could be absent, weakly encoded, represented in directions or interactions that the chosen cosine ranking does not reward strongly enough, or swamped by other variation. What the experiments below measure is narrower and cleaner — whether this scorer reliably orders a selected grade-3 positive above the typed near miss.

When the nearest neighbor is wrong, what kind of wrong is it — and can a similarity score, or a second score, tell the difference?

Two principles this chapter installs

Similarity is not equivalence. A high score means “these occupy nearby regions under this representation and metric,” not “these say the same thing.”

Retrieval is not verification. Finding a passage that mentions a claim is not confirming the claim. Retrieval delivers candidates; deciding what a candidate supports is a separate operation — and, as the demonstration shows, a natural-language-inference score is not automatically that operation either.

A sharper way to say the second principle: a similarity score is not a support verdict. Some near-but-wrong cases are not even false statements — a passage can be true, on-topic, and still fail to support the specific claim the query asks about (partial support, wrong scope, wrong time). The question is support for the requested claim, not metaphysical truth.

Three layers of “does this candidate answer the query?”

It helps to separate three questions that a single cosine number is often asked to answer at once:

  1. Relatedness. Is this candidate about roughly the same thing? Embedding similarity is frequently useful here, and this is the layer retrieval mostly operates on.
  2. Assertion compatibility. Does the candidate preserve the query’s polarity, entity roles, temporal scope, quantities, and claim strength? A generic cosine may not reliably expose this. The near-but-wrong cases in this chapter all live at this layer.
  3. Evidential support. Does this passage actually support the specific claim, under the application’s standard for what counts as support, refutation, or insufficiency? This requires a separate checking operation with access to the claim, an evidence span, and a policy.

Embedding retrieval is primarily a candidate-finding mechanism for layer 1. Its usefulness there still has to be measured on the task. A generic scalar similarity does not by itself settle layer 2, and retrieval alone does not perform the layer-3 support judgment.

A taxonomy of near-but-wrong

“Wrong” is not one thing. Each of these breaks a different semantic dimension that a scalar similarity readout does not reliably separate:

  • Near-paraphrase with changed scope or strength. “The drug reduced symptoms” vs. “The drug reduced symptoms in a subgroup.” Much of the wording and topic are preserved, but the claims are not equivalent.
  • Negation. “X is true” vs. “X is false.” A small surface change can reverse polarity while leaving most lexical and topical content intact (Chapter 1).
  • Lexical / entity trap. A query about “Apple’s 2019 revenue” retrieves a passage dense in “Apple,” “2019,” and “revenue” that is actually about a different metric or a forecast.
  • Same topic, different claim. Both passages are about the treaty; one describes what it proposed, one what was ratified.
  • Relation swap. “A acquired B” vs. “B acquired A.” Same entities, same verb, reversed roles.
  • Temporal mismatch. Correct claim, wrong year: a 2005 population figure returned for a 2024 query.
  • Partial support. The passage supports part of the query and is silent or contradictory on the rest.

The following table is a set of routing hypotheses, not results. The rightmost column marks which types the measured experiment below actually covers.

Failure typeWhat changedWhy a scalar similarity may miss itCandidate next stageMeasured in row 1.7?
Near-paraphrase / changed scopeclaim strength / scopemuch of the surface form and topic is sharedclaim-level decompositionno
Negationpolaritya small edit can flip the assertion while preserving most other contenta scorer validated for polarity, in the right orientationyes (as negation)
Lexical / entity trapthe actual metric or framingquery terms are all presentjoint-encoding reranker, then support checkpartly (as entity-related)
Same topic, different claimwhich proposition is assertedshared topic dominates the scoreclaim decomposition + support checkyes (as topic-related)
Relation swapargument roles / ordersame tokens, same entitiesa scorer that can condition on role/orderyes (as relation-swap)
Temporal mismatchthe time indexthe year may be a low-weight tokenmetadata date filter — only if dates are known, the query has temporal intent, and candidate metadata carries the relevant timeyes (as temporal-mismatch)
Partial supportcoverage of the claimone score cannot express “supports part”claim-level checkingno

The “candidate next stage” column names a separate operation. The demonstration tests exactly one such second stage, and only on the five types RELATE labels.

Demonstration: RELATE typed hard negatives

MEASURED on RELATE v0.1, Wave 1 row 1.7 — artifact experiments/embeddings-from-first-principles/wave1/artifacts/distractor-winrate.json. Bi-encoder all-mpnet-base-v2. The second-stage scorer is an NLI cross-encoder, cross-encoder/nli-deberta-v3-base, which the artifact notes was used as a stand-in for a retrieval cross-encoder not in the local cache.

The experiment is a pairwise discrimination probe, and its exact shape matters for reading the numbers.

For each RELATE query that has a grade-3 positive, the implementation:

  1. selects the first grade-3 positive as the reference positive (p0); a query can have several acceptable positives, so beating p0 does not prove a candidate beats every acceptable answer;
  2. for each labeled hard negative h, grouped by its underlying_relation, scores p0 and h against the query;
  3. counts a bi-encoder distractor win when cos(query, h) > cos(query, p0).

So the reported statistic is the fraction of (query, reference-positive, typed-distractor) triples in which the distractor outscores that one selected positive — by underlying relation type. It is not a top-1 retrieval error rate, not Recall@1, and not a statement that the distractor is the nearest item in the whole corpus. The query set mixes styles — restatements, questions, keyword strings, and longer natural-language queries — and is not filtered. The hard negatives also come from several construction methods: negation, relation-swap, and temporal-mismatch are structured perturbations and provide tighter controls; topic-related and entity-related negatives are matched synthetic distractors that can differ from the reference positive in more than one way. Do not treat all five categories as one-variable minimal pairs.

underlying relation      n      bi-encoder     NLI entailment scorer
entity-related           329      0.040             0.255
negation                 257      0.035             0.339
relation-swap             61      0.164             0.082
temporal-mismatch         24      0.000             0.375
topic-related            266      0.158             0.274

All figures are pairwise distractor-win rates against one selected grade-3 positive.

The bi-encoder’s pairwise outrank rate is modest but strongly type-dependent on this probe. Distractor wins run from 0.164 for relation-swap down to 0.000 for temporal-mismatch. The temporal cell contains only 24 comparisons, so read it as “no distractor won in this sample,” not as immunity to temporal errors. negation and entity-related are also low here (0.035 and 0.040). The important result is the variation by failure type, not a claim that the bi-encoder either “works” or “fails” in aggregate.

The NLI second stage is not a top-k reranker, and it makes things worse for four of the five measured types. For each (query, p0, h) triple, the NLI cross-encoder scores (query, p0) and (query, h) with a three-way softmax whose experiment-side label order is contradiction, entailment, neutral. The candidate score is the entailment probability at index 1; the distractor wins when that probability exceeds the reference positive’s. The input orientation is fixed as (query, candidate). Under the standard premise→hypothesis interpretation of NLI pairs, that direction asks something different from candidate→query support, and many RELATE queries are interrogative or keyword-like rather than declarative hypotheses. The experiment does not isolate which of those mismatches matters. It establishes only the behavior of this NLI-trained pair scorer, with this orientation and score definition — not the behavior of retrieval cross-encoders in general.

With that scorer, relation-swap improves (0.164 → 0.082) and the other four worsen — entity-related 0.040 → 0.255, negation 0.035 → 0.339, temporal-mismatch 0.000 → 0.375, topic-related 0.158 → 0.274. The experiment does not isolate why. Plausible contributors include the pair direction, the mixed and often non-declarative query forms, a distribution mismatch between NLI training data and RELATE, candidate-specificity effects, the NLI model’s objective, and differences among the distractor constructions — as well as genuine polarity and time behavior. What is established is narrower: using this entailment-probability scorer in this orientation increased pairwise distractor outranking for four of the five measured types. The relation-swap improvement is consistent with joint pair encoding giving such a model the capacity to condition on argument order, but this single result does not demonstrate that mechanism.

MEASURED: on RELATE v0.1 typed hard negatives, all-mpnet-base-v2 lets a distractor outscore one selected grade-3 positive in 0–16.4% of pairwise comparisons, with the rate strongly type-dependent. Substituting an NLI entailment-probability score for the cosine, in (query, candidate) orientation, lowered the relation-swap outrank rate (16.4% → 8.2%) and raised it for the other four types (up to 33.9% for negation and 37.5% for temporal-mismatch, n=24).

The lesson is the one the book keeps arriving at from different directions: a second, more elaborate model is not automatically an improvement. Its scoring question must be specified and tested against the distinctions the application needs; a seemingly relevant training objective is not sufficient evidence by itself. Metric ≠ meaning (Chapter 4); detector ≠ policy (Chapter 6); geometry ≠ task quality (Chapter 8); index fidelity ≠ semantic correctness (Chapter 9); and here, second model ≠ validated second stage.

What this chapter establishes and what it does not

Establishes: a taxonomy of near-but-wrong retrieval failures at the assertion-compatibility layer; that on RELATE typed hard negatives a bi-encoder’s pairwise distractor-win rate against one selected positive is modest but type-dependent (0–16.4%); and that an NLI entailment-probability scorer, used pairwise in (query, candidate) orientation, improved one measured type and worsened four.

Does not establish: any full-corpus top-1 retrieval error rate (row 1.7 is pairwise, against p0 only); that these rates hold on other models, corpora, or query distributions; why the NLI scorer worsened four types; that cross-encoders in general fix relation swaps or worsen negation; that a metadata date filter solves temporal mismatch; or how verification should be done. It establishes only that retrieval similarity and this particular pairwise entailment score are not, by themselves, a complete support-verification procedure.

Lab 10: the pairwise distractor probe

MEASURED — artifact experiments/embeddings-from-first-principles/wave1/artifacts/distractor-winrate.json (bi-encoder mpnet-base; NLI cross-encoder cross-encoder/nli-deberta-v3-base as second-stage scorer). REPRODUCIBLE — run_wave1.py 1.7.

Question. Which typed near-misses outscore a known-good answer — under cosine, and under a second-stage score?

Step 1 — define the unit. A triple: (query, selected grade-3 reference positive p0, typed hard negative h).

Step 2 — bi-encoder margin. m_bi = cos(q, p0) − cos(q, h). Distractor wins iff m_bi < 0.

Step 3 — second-stage margin. For the measured NLI scorer: m_nli = P(entailment | q, p0) − P(entailment | q, h), using the entailment index of a softmaxed (query, candidate) prediction. Distractor wins iff m_nli < 0.

Step 4 — aggregate by exact artifact type. Report n, the bi-encoder distractor-win rate, and the NLI distractor-win rate:

Underlying relationnbi-encoderNLI entailment scorer
entity-related3290.0400.255
negation2570.0350.339
relation-swap610.1640.082
temporal-mismatch240.0000.375
topic-related2660.1580.274

Pairwise distractor-win rates against one selected grade-3 positive.

Step 5 — do not stop at the win rate. Record the distribution of m_bi, not just its sign. A positive that wins by 0.30 and a positive that wins by 0.003 both count as “not a distractor win,” and they are very different ranking situations. Chapter 11 takes the next step by making margin against the hardest negative in a defined negative set the central quantity; the current row 1.7 artifact records win rates, not the pairwise margin distribution proposed here.

Step 6 — break down by query style (PROPOSED — no artifact backs this). Row 1.7 mixes restatements, questions, keyword strings, and long natural-language queries. Split the NLI results by query_style; premise→hypothesis entailment scoring on a keyword string or an interrogative is not the same task as on a declarative pair, and the aggregate may be hiding that.

Step 7 — policy comes afterward. A 3.5% negation outrank rate is not automatically “safe to trust.” Whether it is acceptable depends on the cost of a wrong answer, the risk the application tolerates, the coverage you need, latency and cost budgets, and whether a downstream verification stage exists. The probe produces evidence; the action is set elsewhere.

Try it yourself

Build matched triples in your own corpus — one query, one correct passage, one typed near-miss that shares surface form. Score them with your bi-encoder and with your second stage, report per-type win rates and the m_bi distribution rather than one average, and record n for every cell: 0 wins out of 24 and 0 wins out of 10,000 are not the same evidence. Do not reuse the table above as evidence for your own model.

Companion component: the distractor probe

The probe is a measurement procedure. It produces per-type evidence; it does not assign a trust level.

distractor_probe:
  space_id:
  query_set:
  reference_positive_policy:     # e.g. "first grade-3 positive"

  scorer:
    type:                        # bi_encoder | cross_encoder | ...
    model:
    pair_orientation:            # e.g. "(query, candidate)"
    score_definition:            # e.g. "cosine" | "softmax entailment prob, index 1"

  per_type:
    "<relation>":
      n:                         # always stored — a rate without n is not evidence
      distractor_win_rate:
      margin_distribution:       # summary of m = score(p0) - score(h)

    second_stage:
    scorer:                      # same fields as `scorer`
    per_type:
      "<relation>":
        n:
        distractor_win_rate:
        margin_distribution:
        delta_vs_first_stage:

  policy_input:
    measurement_summary:         # evidence carried forward
    required_action:             # trust | rerank | verify — SET ELSEWHERE, per application

The Observatory runs the probe when a model or corpus is registered and attaches the results to the retrieval spec. The chain is probe → evidence → policy decision, not probe → automatic verdict. A measured error rate, per-type deltas, and margin distributions are inputs to a risk decision that the application owns.

Failure modes

  • Reading a pairwise outrank rate as top-1 retrieval error. Tempting because both sound like “the wrong thing won.” Check: row 1.7 compares two candidates against a query, not one candidate against the whole corpus.
  • Treating a similarity score as a support verdict. Tempting because the top result is usually relevant. Check: “near under this representation and metric” is not “supports this claim” — that is a separate operation with separate inputs.
  • Treating a second model as a second opinion. Tempting because more machinery feels safer. Check: a second scorer can move errors in the wrong direction for the task — this NLI scoring rule increased pairwise distractor wins on four of five measured types. Measure the scorer’s actual decision rule and per-type effect before treating it as a safeguard.
  • Using an NLI score without pinning the direction. Tempting because the API returns one number. Check: premise→hypothesis is directional; “query entails candidate” and “candidate supports query” are different questions.
  • Averaging across distractor types. Tempting because one headline number is convenient. Check: the average hides the 0–16% bi-encoder spread and the fact that the second stage reversed which types were easy.
  • Ignoring n. Tempting because a clean 0% looks decisive. Check: 0/24 and 0/10,000 are different strengths of evidence; store n and, where you can, an interval.
  • Letting the probe set policy automatically. Tempting because it would close the loop. Check: risk tolerance is an application decision (Chapter 6’s detector-versus-policy rule again).

What this chapter established

  • A vector scorer can rank an on-topic, lexically similar near-miss above a known-good candidate whenever the distinction turns on polarity, role, time, scope, or claim. Similarity is not equivalence; retrieval is not verification; a similarity score is not a support verdict.
  • The three-layer model separates relatedness, assertion compatibility, and evidential support. Embedding retrieval is a candidate-finding tool for the first; the other two need their own measurements and, for support, a separate checking procedure.
  • Row 1.7 is a pairwise probe against one selected grade-3 positive — not a corpus top-1 error rate. Its bi-encoder distractor-win rate was 0–16.4%, and strongly type-dependent.
  • Substituting an NLI entailment score in (query, candidate) orientation lowered one type’s rate and raised the other four: a second model is not automatically a second opinion. The distractor probe records scorer semantics and per-type n as evidence for a policy decision made elsewhere.

Next

Row 1.7 answers a binary question: did the wrong candidate score higher? But a correct passage that wins by 0.003 and one that wins by 0.40 are in completely different situations, and the win rate cannot tell them apart. Define the margin:

margin = score(correct) − score(hardest negative)

The win rate only checks whether the margin is negative; it discards the magnitude of every win and loss. The next chapter changes the experiment: for each query, compare the correct candidate with the hardest negative in a defined negative set, then track the resulting margin and Recall@1 as the negative set gets harder. A preview from row 1.8: the mean structured-perturbation margin for relation-swap is 0.0252.