Hard Negatives
The same model, the same queries, the same reference positive — only the negative-selection rule changes, and the measured margin moves from 0.47 to 0.06. Define margin precisely, separate selection, geometric, and semantic hardness, and show that the negative set is part of the measurement instrument, not an implementation detail.
Part III — Retrieval Is an Experiment
Same model, four negative sets, four verdicts
MEASURED on RELATE v0.1, Wave 1 row 1.8 — artifact
experiments/embeddings-from-first-principles/wave1/artifacts/margin-collapse.json. Modelall-mpnet-base-v2, unchanged across every row below.
negative-selection rule mean margin selected-positive win rate
same-domain random, up to 5 per query 0.4738 1.0000
in-model selector, as implemented 0.2094 0.9814
structured perturbations 0.0932 0.9331
BM25 lexical top-5 0.0595 0.7621
The model did not change. The 269-query set did not change. The reference-positive rule did not change. Only the rule that picked which passages the positive had to beat changed — and the mean margin ranged from 0.4738 to 0.0595, a nearly eightfold difference between the widest and narrowest measured conditions.
Chapter 10 asked a binary question: did a typed distractor outscore the selected positive? That throws away almost everything the geometry has to say. A positive that beats its hardest competitor by 0.801 to 0.800 and a positive that beats it by 0.801 to 0.301 are both “wins,” and they describe completely different retrieval situations. This chapter installs the quantity that tells them apart — the margin — and then uses it to show something sharper than “hard negatives are harder”: the negative-selection rule is not incidental to a retrieval benchmark. It is part of what the benchmark measures.
What distinction is a benchmark actually testing — and are its negatives hard enough to test it?
The experimental unit and the margin, defined precisely
For each RELATE query with at least one grade-3 positive, the implementation takes the first grade-3 positive as the reference — call it p0. RELATE queries can have more than one acceptable grade-3 positive, so everything below is a statement about p0, not about every acceptable answer. This is the same scope discipline Chapter 10 installed for its distractor probe.
Given a constructed negative set S(q) for query q:
margin(q, S) = score(q, p0) − max_{n ∈ S(q)} score(q, n)
score is cosine on the harness’s L2-normalized vectors. The implementation records a win for the reference positive only when score(q, p0) > max score(q, n) — a strict inequality, so a tie counts as a loss for the positive, not a win. The artifact persists the mean margin per negative set, and a field named recall_at_1_vs_negset. Read that name carefully: it is the fraction of included queries where p0 outscored the hardest candidate in that specific constructed negative set — not full-corpus Recall@1, and not a statement about every item in the index. In this chapter, “selected-positive win rate” is the prose name for that same quantity; the JSON field keeps its own name.
Hardness is relational
A passage is not intrinsically a hard negative. It is hard for a particular query, against a particular reference positive, under a particular scorer. Change any one of those and the same passage can move from “obviously wrong” to “nearly indistinguishable.”
It helps to separate three things that “hard negative” can mean at once:
- Selection hardness — how the candidate was mined: at random, by lexical overlap, by the evaluated model’s own ranking, or by a targeted perturbation.
- Geometric hardness — how small the measured score margin actually is, under the specific model and metric being evaluated.
- Semantic (capability) hardness — which distinction the negative is meant to test: polarity, argument role, time, entity identity, lexical overlap.
These do not line up automatically. A structured relation-swap negative can be semantically sharp — deliberately targeted at one capability — and still turn out geometrically easy for a given model. A BM25-matched negative can be lexically brutal while being semantically unrelated to the query’s actual claim. Measuring the margin is how you find out whether a selection-hard negative is also geometrically hard for this scorer.
How to mine hard negatives
| Mining method | How it selects | Targets | Measured in row 1.8? | Main risk |
|---|---|---|---|---|
| Random (same-domain here) | non-positive items from the target’s own RELATE domain, up to 5 per query | a coarse baseline | yes, as random | can still be topically close by chance |
| In-model (self-adversarial) | the evaluated model’s own ranked non-positives | the model’s own blind spots | yes, as in_model — see the caveat below | label noise; here, also a non-obvious slicing choice |
| Structured perturbation | negate / swap roles / shift dates / weaken quantifiers | specific capabilities (polarity, role, time, quantity) | yes, as structured_perturbation, and per relation | perturbed text can read as unnatural |
| Lexical-overlap matched | top BM25 candidates, excluding labeled positives | the keyword trap | yes, as lexical | high overlap does not guarantee topical relatedness or irrelevance |
| Cross-model | a different model’s similar-but-unlabeled items | avoids letting the evaluated model define its own test | not measured here | inherits the auxiliary model’s own biases |
| Entity-matched | non-positives sharing the query’s named entities | entity confusion | not measured here (it appears in Chapter 10’s taxonomy, not in this artifact) | — |
Good hard-negative sets stratify by type, because an aggregate score can tell you that performance moved while a stratified score can show where the change concentrates. Calling that concentration a capability failure requires the stratum itself to isolate that capability cleanly; a matched or mined negative may differ along several dimensions at once.
The false-negative problem
Mining a negative as “scores high, not labeled positive” does not establish that it is actually irrelevant. An unlabeled high-scoring candidate could be genuinely irrelevant, partially relevant, a valid alternate answer, unlabeled supporting evidence, conflicting evidence worth surfacing, or a near-duplicate of a labeled positive. Treating any of these as a clean negative has different costs depending on where it is used:
- In evaluation, it can count a genuinely good retrieval as an error, deflating a score that was actually correct.
- In training, many contrastive losses will apply pressure that lowers the score of a candidate labeled negative relative to the positive. If that candidate is actually valid support, the training signal points in the wrong direction. The exact geometric effect depends on the loss and parameterization, but the label corruption is concentrated exactly where the benchmark is hardest and the distinction most valuable.
There is no single fix, and “skip the top one or two candidates as likely positives” is not a principled one — a genuine hard negative can sit at rank one, and a missed positive can appear much lower. A more defensible hierarchy:
- Use every known positive and exclusion set you already have to remove obvious contamination.
- Apply provenance or rule-based constraints where the domain supports them (e.g., a structured perturbation’s parent item is known and excludable).
- Adjudicate a sample by hand, or with a higher-quality labeling process than the one that produced the mined set.
- Estimate false-negative or ambiguity prevalence per stratum rather than for the set as a whole — the rate among lexical-matched candidates need not match the rate among structured perturbations.
- Where a candidate cannot be confidently adjudicated, label it ambiguous rather than forcing it into “negative.”
A model can assist step 3, but its judgment is another signal to weigh, not a source of ground truth. The right practice is to estimate and report the false-negative or ambiguity prevalence for the mined set, because it qualifies how the resulting scores should be read — not because it converts into a simple arithmetic correction. Knowing that 8% of a stratum is ambiguous does not tell you, without further assumptions about which queries and ranks are affected, exactly how many points of score are artifact rather than signal.
What hard negatives are for
- Evaluation. They let you probe difficult or targeted confusion regimes that easy candidate sets may hide. Some mined negatives may resemble production traffic; structured perturbations may instead be deliberate stress tests. Either way, the negative distribution must be named. A score against an easy set and a score against a targeted hard set answer different questions even when the model is unchanged.
- Diagnosis. Stratified by mining method and, where available, by targeted capability, hard negatives localize a failure instead of just reporting that a score dropped.
- Training. In many contrastive objectives, once a negative already scores far below the positive it contributes comparatively little useful gradient; a close negative contributes more. That is the mechanism-level reason hard-negative mining can make training more informative — but it depends on the loss function, temperature, batch composition, and current scores, so it is not a universal law about every contrastive objective. And the same closeness that makes a hard negative informative when it is genuine makes a false hard negative especially damaging when it is not: training and evaluation share the false-negative risk, but training additionally has to worry about curriculum, hardness scheduling, and overfitting to whatever model did the mining. One mining recipe built for evaluation is not automatically the right recipe for training.
Demonstration: the negative set is part of the measurement instrument
Walk through how each row in the opening table was actually constructed, because the differences are not what a generic description would suggest.
Random. For each query, up to five non-positive items are drawn from the same RELATE domain as the query’s target item. This is not an arbitrary-corpus random baseline; same-domain sampling already removes the easiest possible confusions. Measured: mean margin 0.4738, win rate 1.0000 — the reference positive separates comfortably from this baseline, on this corpus. That supports “same-domain random negatives are markedly easier to reject than the other three constructions here.” It does not establish that the margin’s size reflects pure topical separation — same-domain random items can themselves share real topical or lexical content, and the artifact does not decompose the margin into causes.
Lexical. The top five BM25-ranked candidates for the query, with all labeled positives removed. Measured: mean margin 0.0595, win rate 0.7621 — the lowest of the four constructions on both statistics. These candidates are not guaranteed to be true negatives or guaranteed to be semantically wrong; they are simply unlabeled-as-positive and lexically prominent. The result establishes that this BM25-selected candidate set was the hardest of the four measured constructions for this scorer under these labels. It does not by itself show that lexical confusability is the model’s weakest capability, because the candidate sets differ in composition and the lexical set was not independently adjudicated for false negatives.
In-model, with a caveat worth stating plainly. The description “the model’s own top-5 non-answers” is not what was run. The implementation ranks all non-positive candidates by the evaluated model’s own score and takes ranks 3 through 7 of that ranking ([2:7] in zero-indexed slicing) — it skips the two highest-scoring non-positive candidates and evaluates against the next five. Measured: mean margin 0.2094, win rate 0.9814. This condition looks easier than the lexical condition, but the experiment does not isolate why. The [2:7] slice certainly means the two highest-scoring non-positive candidates were not tested, so this is not a measurement of true top-5 self-mining. Whether that slice was intended as a false-negative guard, an implementation artifact, or something else cannot be recovered from the result. The defensible conclusion is simply: the condition labeled in_model here is ranks 3–7 after positive exclusion, and its comparatively large margin cannot be used to claim that lexical mining is generally harder than self-mining. A top-5 rerun would be a different experiment; no such rerun is reported here.
Structured perturbation. All hard negatives whose method is structured_perturbation — a count that varies by query, covering negation, relation-swap, and temporal-mismatch. Measured: mean margin 0.0932, win rate 0.9331.
That win rate is worth pausing on. A mean margin of +0.0932 is positive on average. But the win rule is strict: m > 0. A win rate of 0.9331 therefore means roughly 6.7% of included queries had m <= 0 — the reference positive either tied or scored below the hardest selected structured negative. The artifact does not separately report ties, so do not convert every non-win into a strictly negative margin. The broader lesson survives: a positive aggregate mean can coexist with a nonzero tail of query-level non-wins.
Structured margins by relation — a different aggregation. The artifact also reports:
negation 0.1357
relation-swap 0.0252
temporal-mismatch 0.0970
These are not computed the same way as the table above. For the four-way comparison, the margin is per query against the hardest negative in the set. For the by-relation breakdown, the implementation instead takes, for every individual structured hard negative, the pairwise difference score(q, p0) − score(q, that one negative), grouped by its underlying_relation, and averages those pairwise differences. So relation-swap = 0.0252 is a mean pairwise margin against relation-swap perturbations specifically, not “the margin against the single hardest relation-swap negative per query.” Keep the two tables conceptually separate even though they share a corpus and model.
relation-swap has the smallest of the three reported by-relation means — 0.0252, versus negation at 0.1357 and temporal-mismatch at 0.0970. What that supports is narrow and exact: relation-swap has the smallest measured mean pairwise structured margin of the three reported types on this model. It does not establish statistical significance (no standard error, confidence interval, or significance test was computed), operational acceptability (no application-specific margin requirement was defined), or a noise floor (no perturbation/stability baseline was measured). The next question is therefore empirical: how does a 0.0252 mean margin compare with the score variation produced by query paraphrases, model updates, corpus changes, approximate search, or other perturbations that matter to the application? That is the kind of stability and calibration experiment Chapter 14 can motivate; this artifact does not answer it.
How much the margin moved. Random to lexical is the sharpest drop measured: 0.4738 → 0.0595, about 7.96× smaller (roughly an 87% reduction). Random to structured is 0.4738 → 0.0932, about 5.08× smaller (roughly an 80% reduction). Random to in-model is only about 2.26× smaller — a reminder to compare exact ratios per condition rather than reaching for one “5–8×” figure for every non-random set; the in-model condition’s ratio reflects its own [2:7] construction, not a general property of self-mined negatives.
MEASURED: with the model, query set, and reference-positive rule held fixed, changing only the negative-selection rule moved the mean margin from 0.4738 (same-domain random) to 0.0595 (BM25 lexical top-5), through 0.2094 (an
[2:7]-sliced in-model condition) and 0.0932 (structured perturbations). The structured win rate was 0.9331, so about 6.7% of included queries hadm <= 0under the strict win rule; the artifact does not separate ties from negative margins. Within the structured perturbations, the mean pairwise margin against relation-swap negatives, 0.0252, is the smallest of the three reported relation types.
A benchmark is not just a query set, a positive, and a metric. The negative-selection rule is part of the measurement instrument, exactly as the metric and representation are. The observed margin remains a property of the interaction between the model and that benchmark construction; the experiment does not make the model irrelevant. It shows that a retrieval number reported without its candidate competition is missing required context.
What this chapter establishes and what it does not
Establishes: margin is a continuous quantity that a binary win/lose rate discards; on RELATE v0.1 / mpnet-base, four differently constructed negative sets produced mean margins from 0.4738 down to 0.0595 with the model, queries, and reference-positive rule unchanged; the artifact’s recall_at_1_vs_negset means “the reference positive beat the hardest candidate in this constructed set,” not corpus-wide Recall@1; a positive mean margin can coexist with query-level non-wins (for structured perturbations, about 6.7% had m <= 0 under the strict rule, with ties not separately reported); the in_model condition as implemented skips the two top-ranked self-mined candidates and so is not top-5 self-mining; and the by-relation structured margins are pairwise averages, not per-query hardest-negative margins.
Does not establish: that the model is “bad” (it may be excellent at what topical retrieval actually needs); a universal hard-negative recipe; a false-negative rate for any of the four mined sets; statistical significance, a noise floor, or an operational threshold for any margin value; or that calibration fails at these margins — only that small margins motivate measuring stability directly.
Lab 11: build the four negative sets, then look past the mean
MEASURED — artifact
experiments/embeddings-from-first-principles/wave1/artifacts/margin-collapse.json(mpnet-base). REPRODUCIBLE —run_wave1.py 1.8.
Question. How much of an apparent margin is a property of the negatives it was measured against?
Step 1 — freeze the common variables. Record the corpus hash/version, 269-query set, model and revision, query/document embedding protocol, cosine on L2-normalized vectors, and the reference-positive rule (first grade-3 positive). Also record the candidate exclusions per mining method: random, lexical, and in_model explicitly exclude all labeled positives in row 1.8; structured_perturbation instead consumes the query’s designated structured hard-negative records directly.
Step 2 — define each negative set exactly as implemented, not as its name suggests.
random: up to 5 same-domain non-positive items per query.lexical: top-5 BM25 candidates after excluding labeled positives.in_model: the model’s own non-positive ranking, sliced[2:7]— ranks 3–7, not the top 5.structured_perturbation: every hard negative whose method isstructured_perturbation; count varies by query.
Step 3 — compute the per-query hardest-negative margin. m(q, S) = score(q, p0) − max_{n∈S(q)} score(q, n).
Step 4 — aggregate.
| Negative set | Mean margin | Win rate (m > 0) |
|---|---|---|
| random | 0.4738 | 1.0000 |
in_model ([2:7]) | 0.2094 | 0.9814 |
| structured_perturbation | 0.0932 | 0.9331 |
| lexical | 0.0595 | 0.7621 |
Step 5 — stratify the structured perturbations, and note the different aggregation.
negation 0.1357
relation-swap 0.0252
temporal-mismatch 0.0970
These are per-negative pairwise margins averaged within each relation type — not the per-query hardest-negative margin used in Step 4’s table.
Step 6 — inspect the missing distribution (PROPOSED — no artifact backs this). The mean and win rate hide the median, tails, and individual cases. Persist: median margin, low percentiles (p05/p10), counts with m < 0, m = 0, and m > 0, and — where feasible — a bootstrap interval on the mean. Separating ties from inversions matters because the measured win rule treats both as non-wins.
Step 7 — audit false negatives (PROPOSED — no artifact backs this). Sample candidates from each mined stratum and classify each as: true negative, partial support, alternate valid positive, ambiguous, or duplicate/provenance-related to a labeled positive. Report n and the resulting uncertainty per stratum; do not assume every unlabeled candidate is a true negative.
Try it yourself
Mine all four sets on your own queries, using the exact selection rules above (not the
[2:7]quirk unless you mean to reproduce it). Stratify the structured results by perturbation type. Report the ratio between the random-negative margin and each harder condition’s margin — computed exactly, the way this chapter did, rather than as one rounded “several times smaller.” Then run Step 7 on at least 20 candidates from your hardest-scoring stratum before you trust or train on it.
Companion component: the negative-set descriptor
A retrieval margin without this block is missing the information needed to interpret it.
negative_set_descriptor:
corpus_hash:
query_set_id:
reference_positive:
rule: # e.g. "first grade-3 positive"
multiple_positives_possible: true
scorer:
space_id:
metric:
mining:
method: # random | bm25 | in_model | cross_model | structured_perturbation | entity_matched
candidate_pool: # e.g. "same domain", "full corpus"
exclusions: # e.g. "all labeled positives for this query"
count_per_query: # fixed (e.g. 5) or variable — state which
ranking_rule: # e.g. "BM25 desc" | "model score desc"
slice: # e.g. "[0:5]" or "[2:7]" — state the ACTUAL slice used
rng_seed: # row 1.8 code uses the experiment RNG; persist it in the artifact/spec
label_quality:
adjudication_method: # unadjudicated | sampled_human | ...
estimated_false_negative_rate: float | not_measured
ambiguous_rate: float | not_measured
outcomes:
n_queries:
mean_margin:
median_margin: # not_measured for the row-1.8 artifact
fraction_margin_le_zero: # 1 - win_rate, where win_rate is measured
selected_positive_win_rate: # the artifact's `recall_at_1_vs_negset`
strata:
"<relation or type>":
n:
mean_pairwise_margin_or_mean_margin: # state which aggregation was used
For the row-1.8 artifact specifically: mining.slice for in_model is [2:7], not [0:5]; outcomes.median_margin is not measured and should be recorded as such rather than left implicitly equal to the mean; and the strata entries under structured_perturbation use the pairwise aggregation, which the descriptor should say explicitly so it is never confused with the top-level outcomes.mean_margin aggregation.
For margin-style evaluations, the Observatory will not surface the number without an attached negative-set descriptor of this kind. Other retrieval metrics need the corresponding evaluation context instead: for nDCG over a graded corpus, that means at least the candidate pool and relevance definition rather than a separately mined negative set. The general rule is broader than this chapter’s schema: no retrieval metric is displayed without the information that defines its comparison set and labels. A margin without its negative-selection rule is an orphaned measurement.
Failure modes
- Reporting a retrieval score without its negative distribution. Tempting because one number is easy to share. Check: the same model produced margins from 0.06 to 0.47 here on the same queries — ask what competed against the positive.
- Treating an unlabeled candidate as a known negative. Tempting because “not labeled positive” reads as “negative.” Check: it can be a genuine negative, a missed positive, partial support, or ambiguous — adjudicate before training or scoring on it as ground truth.
- Calling the
in_modelcondition “top-5.” Tempting because that is what self-mining usually means. Check: the measured implementation used ranks 3–7 ([2:7]); its margin is not evidence about true top-5 self-mining. - Comparing mean margins across methods without checking their provenance. Tempting because they are all “margins.” Check: different pool sizes, exclusion rules, and slices are answering different questions; state the mining rule before comparing numbers.
- Trusting the mean margin alone. Tempting because it is one clean number. Check: structured perturbations averaged +0.0932 while about 6.7% of included queries failed the strict
m > 0win condition (m <= 0; ties were not separately reported) — the mean hid a real non-win tail. - Claiming significance without an interval. Tempting because 0.0252 looks small enough to matter. Check: no standard error, confidence interval, or significance test was computed here; say “the measured mean” and stop there.
- Training directly on mined negatives without a false-negative check. Tempting because mining is automatic and cheap. Check: a contrastive objective will push the model away from any false negative in the batch — sample and adjudicate before training on a new mining rule.
What this chapter established
- Margin —
score(reference positive) − score(hardest negative in the set)— is the continuous quantity a binary win/lose rate throws away. - With the model, queries, and reference-positive rule held fixed, changing only the negative-selection rule moved the mean margin from 0.4738 to 0.0595. That belongs to the model × benchmark-construction interaction, not to either one alone: the negative set is part of the measurement instrument, exactly as the metric and the representation are.
- Hardness is relational — selection hardness (how a negative was mined), geometric hardness (how small its margin turned out), and semantic hardness (which distinction it targets) are distinct. The
in_modelcondition’s actual[2:7]slice is why provenance has to be read before any of its margin is interpreted. - A positive mean margin coexists with a real non-win tail (~6.7% of queries under the structured rule), and the by-relation margins use a different aggregation from the main table. Neither supports a statistical, operational, or noise-floor claim.
- False negatives corrupt labels precisely at the boundary hard-negative mining exists to probe. The negative-set descriptor therefore travels with every margin — rule, candidate pool, exclusions, scorer, label-quality status — and records
not_measuredrather than implying an estimate exists.
Next
A retrieval system does not consist of one isolated metric. By now the book has accumulated a representation, a score function, a candidate-eligibility rule, an exact-or-approximate execution strategy, a taxonomy of typed distractors, and a negative-selection rule with its own margin. The score never travels alone — it arrives wrapped in a candidate set, a mining rule, a cutoff, a ranking policy, and a budget. The next chapter assembles those decisions into an explicit chain and asks how changing one link changes what the rest of the system receives.