← Embeddings From First Principles

How Do You Evaluate an Embedding?

An nDCG score is not a property of a model. It belongs to a query population, a candidate universe, a relevance definition, a metric, an aggregation rule, and a model-use protocol. Measure the same three encoders on RELATE v0.1 and v0.2 — same item pool, same metric implementation — and watch the aggregate winner change when the query workload expands.

Part IV — Measuring the Representation

Same encoders, same item pool, different winner

Three encoders. One item pool. One typed-pair pool. One metric implementation. Three relevance definitions, applied consistently. The encoder models and item-side corpus are unchanged between the two releases; what expands is the query population, from 269 queries to 411.

RELATE v0.1 (269 queries):    all-mpnet-base-v2 wins every relevance definition
RELATE v0.2 (411 queries):    bge-large-en-v1.5 wins every relevance definition

That is not a hypothetical. It is Wave 1 row 1.9, run twice against two frozen releases, and it is this chapter’s demonstration in miniature before the demonstration itself. The models did not change. The evaluation workload did. The winner changed.

Chapter 12 closed Part III by making the retrieval system explicit — representation, candidate eligibility, scoring and fusion, selection, execution, context assembly, and the evaluation definition that judges the result. It ended on a deliberate move: once that surrounding policy is written down and held fixed, a cleaner question becomes askable — with everything around the representation held still, how good is the representation itself? This chapter opens Part IV by taking that question at face value, and the RELATE result above is exactly why it cannot be answered with a single number.

Before formalizing that, one illustrative scenario, worth keeping because the intuition it captures is real even though the specific numbers are invented for illustration: a team picks the top model on a public embedding leaderboard, screened across many tasks. In their product — retrieval over dense technical documentation, heavy on entity names and version numbers — it underperforms a smaller, older model. Nothing was misconfigured. The leaderboard measured performance under a population of tasks and a weighting policy someone else chose; the product is one task, evaluated under a relevance definition, query population, and candidate universe the leaderboard never saw. This is a scenario to build intuition, not a measurement this book made — the RELATE result above is the measurement, and it makes the same point with numbers you can trace to a file.

What does an embedding evaluation number measure, and under what conditions can two representations be compared by it?

The evaluation contract

“The best embedding model” is not yet a question, because it omits everything that would make it answerable. Best at what task, judged by what relevance definition, over what candidate pool, asked by what query population, scored by what metric, aggregated how, and using each model under whose declared interface? Change any one of those and the ranking can change without touching a single model weight — which is precisely what the opening result just showed by changing exactly one of them.

Write the dependency out:

evaluation result
  =
  representation
  ×  query population
  ×  candidate universe
  ×  relevance definition
  ×  metric (+ k, gain rule, aggregation)
  ×  model-use protocol

An nDCG@10 = 0.88 is not a free-standing property of “the model.” It is an observation produced under a specified evaluation contract — the same discipline Chapter 11 installed for a margin, which is incomplete evidence without the negative-selection rule that produced it, and the one Chapter 12 installed for a retrieval result, which is incomplete without the policy and provenance behind it. This chapter’s version of that rule:

A score is never stored without enough information to reconstruct what it means.

A compact way to hold four of these terms in mind at once, because they answer four genuinely different questions:

Metric says what earns credit. Relevance definition says what counts. Query population says what gets asked. Candidate universe says what can compete.

The retrieval metrics, and what each one keeps

Given a ranked list of candidates with relevance labels, every retrieval metric is a rule for compressing that list into one number — and every such rule throws information away on purpose. Knowing what a metric keeps and what it discards is more useful than memorizing a formula.

  • Precision@k. Fraction of the top k that are relevant. It knows nothing about relevant items outside the top k, and nothing about their order within it.
  • Recall@k. Fraction of all relevant items that made the top k. It rewards coverage and does not penalize a top k cluttered with irrelevant items, as long as the relevant ones are somewhere in it.
  • Reciprocal Rank (RR). 1 / rank of the first relevant item, computed per query. It cares about exactly one thing: how far down the list a user or a model has to look before finding something useful, and it stops caring the instant that item is found — a second and third relevant item contribute nothing.
  • MRR. The mean of RR across queries. RR and MRR are not the same object: one is a per-query number, the other is an aggregate over a query population, and reporting “MRR = 0.7” for one query is a category error.
  • nDCG@k. Discounted cumulative gain, normalized against the best possible ordering of the same relevance labels. It handles graded relevance (not just relevant/irrelevant) and discounts by rank position, so a relevant item at position 1 counts more than the same item at position 9. In this book’s own implementation, the gain and discount are exponential and logarithmic respectively:
DCG@k  = Σ_{i=1}^{k}  (2^grade_i − 1) / log2(i + 1)
nDCG@k = DCG@k / ideal_DCG@k

grade_i is the relevance label of the item at rank i; ideal_DCG@k is DCG@k computed on the same set of labels sorted in the best possible order. The 2^grade − 1 gain means a grade-3 item is worth far more than a grade-1 item, not merely three times as much — a detail worth knowing before comparing nDCG scores computed under a different grading scale.

  • MAP. Mean average precision — average precision-at-each-relevant-rank within one query, then the mean of that across queries. It rewards packing every relevant item early throughout the list, not just the first one.
  • Rank correlation (Spearman, Kendall). For semantic-textual-similarity-style evaluation, this compares the model’s own similarity ordering over labeled pairs against a human-judged ordering. It is a different object from retrieval nDCG: there is no ranked list of candidates being retrieved, only a set of pairs whose relative similarity ordering is being checked. Its output also ignores absolute score calibration entirely — two scorers that rank pairs identically can disagree completely about what a “high” score is (Chapter 14’s subject).
MetricWhat it keepsWhat it discardsWhen it is informative
Precision@krelevance density inside a fixed slot budgetanything below rank k, and order within the top kthe consumer only ever sees k results and scans all of them
Recall@kwhether every relevant item made the cuthow cluttered the top k is with noisecoverage matters more than order — an exhaustive-retrieval consumer
RR (per query) / MRR (mean)how far down the list to the first useful itemevery relevant item after the firstthe consumer stops at the first hit
nDCG@kgraded relevance and position, both at onceabsolute score magnitude; ties beyond the gain function’s resolutionrelevance is graded, not binary, and position matters throughout the list
MAPprecision at every relevant rank, then averagedfine distinctions once every relevant item is retrieved somewherethe consumer wants a fully ordered, exhaustively-relevant list
Spearman / Kendallwhether the ordering of similarity judgments agrees with a referencethe actual magnitude of any scorejudging a similarity function against human pairwise judgments, not a retrieval task

The teaching point is not “memorize six formulas.” It is that each metric answers a different question about the same ranked list, and picking one is already a decision about what counts as success. The principle that should replace any fixed prescription: choose the metric whose incentives match how the ranked list is actually consumed, then validate that proxy against downstream application quality — because every one of these remains a proxy. A system that reads only the first retrieved passage is well served by something RR-like, but MRR still rewards moving a relevant item from rank 10 to rank 2 even when only rank 1 is ever consumed; a system fed several passages cares more about recall, precision within the fed set, and how graded relevance is distributed across them. Neither choice, by itself, measures whether the downstream generator used the passage correctly, ignored it, or was confused by two conflicting ones — that gap is exactly why Chapters 10 and 12 kept retrieval and verification as separate stages, and why this chapter’s metric is a representation-level proxy, not a final verdict on the application.

Representation-level and application-level evaluation

Chapter 12 made the surrounding retrieval policy explicit precisely so a narrower question could be asked without pretending it is context-free. For a representation-level comparison, freeze the parts of the experiment that define the task — candidate corpus, query population, relevance definition, metric/readout, and comparison procedure — then vary the representation deliberately. The representation protocol itself is recorded per model, because fair comparison may require different valid query/document prefixes, roles, or instructions.

  • Representation-level evaluation. Performance of a representation under a deliberately fixed task and comparison contract — one metric, one candidate pool, one relevance definition, one query population — while each model is invoked through its declared valid representation protocol.
  • Application-level evaluation. Performance of the full deployed system: the representation, wrapped in the retrieval policy, feeding a downstream consumer, under real latency, cost, and error-cost constraints.

There is no task-free scalar that deserves the name “representation quality.” A representation useful for retrieval, clustering, classification-by-embedding, semantic-textual-similarity, and entity matching is being asked five different questions, and nothing guarantees the same model answers all five equally well. What earlier chapters called representation quality is better read as representation-level evaluation under a stated readout — which is exactly what this chapter’s demonstration measures, and exactly why it can only speak for the readout it used.

The gap between representation-level and application-level results tends to open along a few recurring seams:

  • Domain shift. A benchmark may underrepresent the domains that matter to an application — legal, biomedical, code, logs, or something narrower — so transfer has to be measured rather than assumed.
  • Relevance definition. “Relevant” on a benchmark may mean same topic; yours may mean answers this exact question, or same entity and same year — the three-layer distinction Chapter 10 installed (relatedness, assertion compatibility, evidential support) is exactly what a relevance definition has to choose among.
  • Query style. The benchmark and the application may ask differently: one may use well-formed questions while the other contains keyword fragments, pasted errors, long requests, or another distribution entirely. The mismatch, not any one style, is the risk.
  • Consumer. A human scanning ten results tolerates noise a model reading one result does not.
  • Model-use protocol. Query and document roles, normalization, prefixes, instruction strings, and truncation limits are part of a model’s declared interface — evaluating a model outside that interface measures model plus misuse, not the representation the model card describes. Two asymmetric encoders forced through identical, undifferentiated preprocessing are not being compared fairly even if every other setting matches.

A useful public benchmark provides real comparative evidence — under its own evaluation distribution. Whether that evidence transfers to a different domain, relevance definition, query population, or model-use protocol is an empirical question, not something the benchmark’s popularity settles for you.

Benchmark hygiene as measurement hygiene

None of the following is a leaderboard-criticism checklist. Each is a specific way an evaluation contract can go unstated, and an unstated contract is exactly what makes a comparison misleading.

  • Workload mismatch. Benchmark queries differ from the application’s queries in style, length, and difficulty — the RELATE demonstration below is this failure mode, deliberately constructed and measured.
  • Relevance mismatch. The benchmark’s labels reward a different notion of relevance than the one the application needs.
  • Metric mismatch. The scoring rule rewards ranking behavior the actual consumer does not use.
  • Aggregate masking. A benchmark average is a weighting policy over many evaluations — which tasks are included, how many examples each contributes, whether scores are macro- or micro-averaged, which metric each task uses. A model averaging well across many tasks can still be mediocre on the one task that matters, and the average alone cannot tell you that; read the per-task breakdown for the tasks that resemble yours.
  • Exposure / contamination risk. Public, widely circulated benchmarks carry a real risk that some of their material appeared in a model’s training data, because training provenance is frequently incomplete or undisclosed. That is a risk to weigh, not a fact to assert about a specific model-benchmark pair without evidence — “this benchmark is public, therefore this score is contaminated” is exactly as unsupported as “this eval is private, therefore this score is trustworthy.”
  • Development overfitting. A private evaluation set, repeatedly used to choose among candidate models or tune a pipeline, gradually becomes a development set. Its results then estimate how well a system was tuned against that set, not how well it will generalize past it.
  • Weak discriminative power. Under an easy workload, scores can cluster in a narrow high-scoring range, leaving little room to separate systems confidently. This is a real, checkable property of an evaluation — narrow score spread — and it is a distinct claim from “the difference is not statistically significant,” which requires an uncertainty estimate this book has not computed for the demonstration below. Do not use one phrase for the other.

These risks motivate a ladder rather than a single verdict on any one evaluation:

public benchmark              → coarse screening: eliminate obviously weak candidates,
                                 identify model families, check against a shared task distribution
      ↓
application-like development
eval                          → engineering comparison: representative queries, your
                                 relevance definition, used to iterate on the pipeline
      ↓
sealed / holdout eval         → final model-selection evidence: not touched during
                                 iteration, so its result is not already optimized against

None of these stages makes the others unnecessary, and none of them is bureaucracy for its own sake — each answers a question the one before it cannot. “Private” is not a synonym for “valid”: a private set can still be too small, too easy, badly labeled, unrepresentative of the real query distribution, dominated by one query type, or itself repeatedly overfit through iterative model selection until it stops functioning as evidence and starts functioning as a target. The property that actually earns trust is not privacy. It is provenance, distribution, and holdout discipline that match the decision being made — and privacy is one common, practical way of reducing public-benchmark exposure risk, not a guarantee of validity on its own.

Demonstration: RELATE v0.1 and v0.2

MEASURED on RELATE, Wave 1 row 1.9 — artifacts experiments/embeddings-from-first-principles/wave1/artifacts/relevance-definition-sweep.json (v0.1) and .../v02/relevance-definition-sweep.json (v0.2). Three encoders, item pool and typed-pair pool held fixed, three relevance definitions, nDCG@10.

A word on what changes between releases and what does not, because getting this wrong is the easiest way to misread the table. RELATE v0.2 is not a 142-query evaluation. It is v0.1’s 269 queries — described in the release documentation as mostly near-restatements of their answer sentence, not exclusively so — plus 142 newly constructed harder queries: indirect phrasing with a genuine lexical and semantic gap to the answer (role disambiguation, claim-strength qualification, temporal qualification, indirect reference), each with roughly five competing distractors rather than v0.1’s roughly 2.5. The item pool and the typed-pair pool are byte-identical between the two releases — confirmed by an exact file hash match on items.jsonl and pairs.jsonl — and only queries.jsonl grows, from 269 rows to 411. Row 1.9’s implementation loops for q in queries over whatever release is loaded, so the v0.2 numbers below are averaged over all 411 queries, not over the 142 additions in isolation. No artifact in this book reports a hard-queries-only score, and this chapter does not invent one.

                              v0.1 — 269 queries                v0.2 — 411 queries
                                                          (269 original + 142 harder additions)

relevance definition     MiniLM   mpnet    bge-large      MiniLM   mpnet    bge-large
answers_only (>=3)       0.9431   0.9485   0.9330          0.8383   0.8466   0.8522
supports (>=2)           0.9362   0.9523   0.9385          0.8514   0.8750   0.8805
on_topic (>=1)           0.9357   0.9518   0.9370          0.8511   0.8747   0.8796

winner, every definition     all-mpnet-base-v2                bge-large-en-v1.5

Read this in two separate passes, because it answers two separate questions.

Pass A — the evaluation workload changed. Model, item pool, typed-pair pool, metric implementation, and all three relevance definitions are identical across the two blocks; only the query population grew by 142 harder items. Two things happened at once, and they are not the same observation: every model’s aggregate nDCG@10 dropped — from the 0.93–0.95 band down to 0.84–0.88 — and the model ordering changed — from mpnet-base winning every definition to bge-large winning every definition. The first fact says the enlarged workload is harder for all three models under this metric. The second says the model ranking is conditional on that workload. Together they support a specific, bounded claim: adding a population of harder queries was sufficient to change which representation the evaluation calls best. They do not support “bge-large is universally better on hard queries” — this is one synthetic corpus, three models, no query-resampling estimate, and no evidence about any other hard-query distribution.

It is tempting to explain the v0.1 result by saying the query set — not the models — was “the ceiling,” or that v0.1 was “saturated,” or that “nothing separates the models” on v0.1. Resist all three. Nothing here measures a formal ceiling or establishes that the models contributed nothing to the clustering; the models do have different measured scores on v0.1, just closer together than on v0.2. The defensible statement is narrower and still useful: under the v0.1 protocol, these three models occupied a narrow, high-scoring range with limited numerical separation; expanding the query workload with intentionally harder items changed both the score level and the winner. Whether the v0.1 differences are statistically distinguishable is a separate question because this experiment reports no resampling or uncertainty estimate.

Pass B — the relevance threshold did not change the winner. Within v0.1, mpnet-base wins at >=3, >=2, and >=1 alike. Within v0.2, bge-large wins at all three thresholds alike. The query-population change reordered the models; the tested relevance definitions did not. This is a genuine negative result, and it is worth stating plainly rather than explaining away. It would be easy to reach for “the three definitions are highly correlated on this corpus” as the reason — but no correlation statistic was computed, and that explanation is an unverified guess wearing the clothes of a finding. What is actually established: on these two frozen releases, moving the relevance bar from “answers the query” down to “merely on-topic” did not change which model came out ahead. Whether a corpus with a different mix of grade-1, grade-2, and grade-3 material would behave differently is a real question this experiment does not answer — it is a hypothesis for the next release to test, not a result this one produced.

One more provenance detail, small but worth naming because the book has taught you to check exactly this kind of thing: the v0.2 artifact’s provenance.note field still reads “RELATE v0.1 is a synthetic controlled probe” and still points at the v0.1 datasheet, while the artifact lives under the v0.2 output path and carries the different corpus hash produced by that release. The conflict is not something prose should silently normalize. Bind the hash to an explicit release identifier and the exact query-set artifact, so provenance can be checked mechanically rather than inferred from a copied note. The evaluation card below does exactly that.

What this chapter establishes and what it does not

Establishes: the standard retrieval and similarity metrics, what each keeps and discards, and that metric choice is a decision about which ranking behavior earns credit; the representation-level / application-level distinction, with representation-level evaluation defined as a fixed-protocol comparison rather than a task-free property; benchmark hygiene as a set of specific, checkable risks rather than a blanket dismissal of public benchmarks; and, measured, that on RELATE the query-population change from v0.1 (269 queries) to v0.2 (411 queries, item and pair pools unchanged) was sufficient to move every model’s aggregate nDCG@10 down and to flip the winning model from mpnet-base to bge-large under all three tested relevance definitions, while the relevance definitions themselves did not reorder the winner within either release.

Does not establish: why relevance-threshold invariance held on this corpus (no correlation between the definitions was measured); that bge-large is universally superior on hard queries, or on any query population outside this one; a corpus-independent claim that harder queries always separate models or that easier queries never do; any confidence interval, bootstrap distribution, or significance test over the query sample (none was computed); a hard-queries-only score for the 142 additions (none exists in the reported artifact); or that public benchmarks are without value — they remain useful for coarse screening even though they do not settle an application-level comparison.

Lab 13: build and audit an evaluation contract

MEASURED — artifacts wave1/artifacts/relevance-definition-sweep.json (v0.1) and wave1/artifacts/v02/relevance-definition-sweep.json (v0.2). REPRODUCIBLE — python run_wave1.py 1.9 for v0.1; RELATE_RELEASE=relate-0.2.0 python run_wave1.py 1.9 for v0.2.

Question. Does your evaluation actually separate the candidates you are choosing among — and can you say exactly what experiment produced each number?

Step 1 — freeze the evaluation contract before running anything. Candidate corpus, version, and hash; query-set version; relevance-definition version; metric, k, and aggregation rule; the query/document representation protocol for every model under test; code hash; declared slices; and whether this run is development or sealed-holdout.

Step 2 — reproduce v0.1. 269 queries, three models, three relevance definitions. Record the exact values above; do not round them away.

Step 3 — reproduce v0.2 correctly. RELATE_RELEASE=relate-0.2.0 python run_wave1.py 1.9. State plainly in your own notes: v0.2 contains 411 total queries, not 142 — the 142 are additions to, not a replacement for, v0.1’s query set.

Step 4 — separate the two effects. For each model, how much did its score move between releases? Separately: did the model ordering change? These are different questions with different answers here — every model’s score dropped, and only the ordering among the three flipped.

Step 5 — sweep the relevance definition within each release. Confirm the winner is unchanged across >=3, >=2, >=1 on both v0.1 and v0.2. Record this as a negative result, not a gap to explain away.

Step 6 — audit the artifact’s own provenance. Check whether the note describing the corpus matches the hash actually used. On the v0.2 artifact it does not; record the release identifier explicitly in your own evaluation card rather than trusting a free-text field.

Step 7 (PROPOSED — no artifact backs this) — estimate uncertainty. Bootstrap the query set with replacement, recompute the winner on each resample, and report how often each model wins. Or hold out a second, sealed query batch and check whether the same model wins on it. Neither has been run for this chapter; do not report a confidence interval or a “statistically significant” claim without doing one of these first.

Step 8 — build an application evaluation. Where possible, use real or application-representative queries rather than another synthetic probe. There is no universal sufficient sample size — the number of queries you need depends on how variable your scores are, how many slices you need to cover, and how small a difference you need to detect reliably. A practical starting range for a first pass is 100–300 queries; treat that as a pragmatic starting point to refine with your own variance, not as a statistically sufficient sample on its own.

Try it yourself

Collect real or application-representative queries from your own logs or usage, write your relevance definition down in words before you look at any scores, evaluate three to five candidate models under your metric plus one alternative, and record operational measurements — p95 latency, cost, vector dimension, max sequence length — alongside the quality scores rather than folding them into one number. Break the result down by query type or another slice that matters to you. If a public leaderboard would have picked differently, do not automatically defer to either ranking. Ask which one actually shares your query distribution, candidate universe, relevance definition, and metric — that evaluation, not the more prestigious one, is the more relevant evidence for your decision.

Companion component: the evaluation card

Chapter 12’s artifact made the retrieval system explicit — policy, provenance, execution, trace, and evaluation kept as separate objects. This chapter’s artifact does the same for a comparison: the contract that defines what is being measured, kept separate from the result any one representation produced under it.

evaluation_card:
  id / version:
  candidate_pool:
    corpus_id:
    corpus_hash:
  queries:
    query_set_id:
    n:
    source:               <public | private | synthetic-probe>
    release:              <explicit identifier — never inferred from a free-text note>
    development_or_holdout:
  relevance:
    definition_id:
    written_definition:   <not "relevant" — the actual threshold or rule>
    label_source:
  metric:
    name:
    k:
    gain_rule:
    aggregation:
    consumer_rationale:   <why this metric matches how results are consumed>
  representation_protocol_policy:
    requirement:          <use each model's declared valid query/document protocol>
    protocol_fields:      [query_role, document_role, normalization, prefix_or_instruction, truncation]
  slices:                 [query_type, domain, entity-heavy, temporal, lexical-gap, ...]
  provenance:
    code_hash:
    artifact_path:
    known_metadata_issues: <e.g. "note field copied from a prior release">

evaluation_observation:
  evaluation_card_id:
  space_record_id:        <from Ch 1/17>
  model_revision:
  resolved_representation_protocol:
    query_role:
    document_role:
    normalization:
    prefix_or_instruction:
    truncation:
  overall_score:
  per_slice_scores:
  operational:
    latency_p95:
    throughput:
    memory:
    cost:
    vector_dim:
    max_sequence_length:

The split is the point: evaluation_card defines the experiment; evaluation_observation records what one representation scored under it. Evaluation definition and the result produced under it are two different objects, and neither is complete evidence without the other — exactly Chapter 11’s margin-and-negative-set discipline, carried one layer up.

Two fields deserve emphasis because they are the ones a quick comparison tends to skip. representation_protocol_policy belongs to the shared card because it defines the fairness rule for the comparison; resolved_representation_protocol belongs to each observation because different models may satisfy that rule with different valid prefixes, query/document roles, normalization, or truncation. Forcing every model through the same literal preprocessing can measure model plus misuse rather than the representation under its declared interface. And operational stays separate from the quality fields on purpose: latency, throughput, memory, cost, dimension, and sequence limits answer a different question than nDCG does. Quality evaluation produces evidence; the trade-off between quality and operating cost is a policy choice for the application — Chapter 6’s detector-versus-policy separation, again.

Embedding Observatory behavior

An Observatory that stores leaderboard numbers is not doing more than a spreadsheet. What it should refuse to do is silently rank two evaluation_observations when their task contracts disagree — different query sets, relevance definitions, candidate universes, metrics, aggregation rules, or protocol requirements. Model-specific resolved preprocessing may legitimately differ when two encoders have different declared interfaces; the Observatory should verify that each observation satisfied the same protocol policy, not demand byte-identical prefixes or roles. Chapter 12 established that a metric must name the stage or output it actually observed. This chapter adds the analogous rule one layer up: a model score must name the evaluation contract that produced it. Two numbers called nDCG@10 are comparable only after those contracts are shown to define the same comparison.

Failure modes

  • Picking by leaderboard average. Tempting because it is a single, already-computed number. Check: the average is a weighting policy over someone else’s task mixture; read the per-task breakdown for tasks that resemble yours.
  • Treating “private” as “valid.” Tempting because privacy feels like it rules out contamination. Check: a private eval can still be unrepresentative, too small, badly labeled, or repeatedly overfit through iterative model selection until it stops being holdout evidence.
  • Leaving relevance undefined. Tempting because “relevant” feels self-explanatory. Check: write the threshold down; different labelers, and different models, will otherwise be scored against different implicit targets.
  • Metric–consumer mismatch. Tempting because one metric is the field’s default. Check: ask how the ranked list is actually consumed before choosing what earns credit.
  • Comparing scores from different evaluation cards. Tempting because both numbers say nDCG@10. Check: the query population, candidate pool, relevance definition, and model-use protocol all have to match before the comparison means anything.
  • Ignoring model-use protocol. Tempting because uniform preprocessing feels fairer. Check: forcing an asymmetric encoder through symmetric preprocessing tests a misuse of the model, not the representation.
  • Reporting only the aggregate. Tempting because one number is easy to headline. Check: the winner on average can lose on the slice that actually matters to the application.
  • Calling a narrow score gap “significant” or “within noise.” Tempting because both phrases sound rigorous. Check: neither is supported without a computed interval or resampling result — this chapter’s own RELATE tables carry neither.
  • Trusting a stale provenance note. Tempting because the note is right there in the artifact. Check: confirm the release identifier and corpus hash the run actually used — the v0.2 artifact’s own note is the concrete reminder that free text drifts and hashes do not.

What this chapter established

  • An evaluation result is a function of representation, query population, candidate universe, relevance definition, metric (with its k, gain rule, and aggregation), and model-use protocol — never a property of the model alone.
  • Precision@k, Recall@k, RR/MRR, nDCG@k, MAP, and rank correlation each keep and discard different information. Metric choice should match how results are consumed and then be validated against downstream quality, never assumed sufficient on its own — and “representation quality” as a task-free scalar does not exist.
  • Measured: changing only the query population flipped the winner. Moving from 269 queries to 411 — same items, same typed pairs, byte-identical pools — reordered the models from mpnet-base to bge-large under all three relevance definitions, while sweeping the relevance definition within either release changed nothing at all.
  • The evaluation card separates the experiment’s definition from any one representation’s result under it, carries the policy for valid model use, and lets the Observatory refuse to rank results whose task contracts are incompatible.

Next

Every score in this chapter answered an ordering question: did items receiving more credit under the stated relevance definition land in better positions? That is what nDCG, Recall@k, and MRR can tell you, and it is all they can tell you. None of them says what a single raw similarity score should cause a system to do. A cosine of 0.81 between a query and a candidate is not self-interpreting — accept, reject, escalate, duplicate, or no action depends on the score distributions for task-defined positive and negative cases, plus the cost of each error. Those distributions can overlap even when a ranking metric looks healthy. Chapter 13 tells you how to judge an ordering under an evaluation contract. Chapter 14 asks whether the raw scores behind that ordering can support a calibrated decision.