← Embeddings From First Principles

Did the Bridge Preserve the Space?

A preservation metric measures one property, not bridge quality. Audit the MiniLM-to-mpnet Procrustes bridge's eight measured fields against their actual implementations, discover that a reassuring 0.86 relation-profile correlation hides a paraphrase-negation contrast flipping sign, and learn why an artifact's filename is not evidence — the supervised-bridge-ceiling experiment never leaves source space at all.

Part VI — Crossing Embedding Spaces · What did each measurement actually certify?

One bridge, eight measurements

Chapter 20 built the machinery that turns preservation evidence into a scoped authorization — a requirement, a measured value, and a decision, all cited together. What it deliberately did not do is ask how much any one of those measured values is actually entitled to say. That is this chapter’s entire job.

A preservation metric is a sensor. Before trusting its reading, ask which failure modes the sensor is even capable of seeing.

Continue with the same bridge Chapter 20 used as its running example — all-MiniLM-L6-v2 (384-d) → all-mpnet-base-v2 (768-d), Procrustes, fitted on split_entity:train anchors and evaluated on held-out split_entity:test entity families. Its full measured profile from Wave 3 row 3.4:

coordinate reconstruction                0.4658
10-NN neighborhood overlap               0.7434
retrieval nDCG@10 ratio                  0.9148
rank-triplet agreement                   0.7098
calibration-transfer score               0.7566
relation-profile Pearson correlation     0.8605
structured hard-negative margin ratio    0.4735
held-out/train reconstruction ratio      0.5562

Eight numbers — not eight verdicts. Chapter 20 already established why: a measurement is evidence a policy can cite, never a decision by itself. This chapter’s task is narrower and comes first — before any policy gets to cite these numbers, each one has to be understood precisely enough that citing it means something. Downstream consumers do not read “the bridge.” They read specific, different functions of the translated vectors: a retrieval system reads which items rank highest; a ranked UI reads their order; a router reads their grouping; a thresholded filter reads their absolute score against a fixed cutoff; a relation-sensitive consumer reads the margin between two specific relation types. A single number cannot answer all of those questions, because they are not the same question.

Which preservation metric answers which question — and what does a good score on one metric fail to tell you about the others?

Generic questions versus the measured instrument

Before auditing what Wave 3 actually measured, it is worth naming the wider space of questions a bridge could in principle be evaluated against, and being honest about which of them row 3.4 answers and which it does not. This is the same discipline Chapter 15 applied to conceptual diagnostics versus the six features row 1.12 actually measured, and Chapter 19 applied to CCA as a real method that nonetheless was not run.

Generic preservation questions one could ask of any bridge: paired-target reconstruction (cosine or MSE to the corresponding native vector); counterpart recovery (Recall@1, Recall@10, MRR — does the translated vector’s own paired target actually turn up nearby?); neighborhood overlap; rank correlation (Kendall’s τ, Spearman); task retrieval quality (Recall@k, MRR, nDCG@k); clustering agreement (ARI, NMI, V-measure); calibration transfer; relation-specific behavior; hard-negative margins; round-trip stability.

What Wave 3 row 3.4 actually measured, for every bridge in this book so far: eight scalar preservation summaries — the fields listed above — plus per-relation native and bridged mean-cosine detail that row 3.7 later extracts. It did not compute MSE — coordinate_reconstruction is mean cosine. It did not compute paired-target Recall@1 or Recall@10 — that is a genuinely different question, and Chapter 22 measures it directly on a larger benchmark. It did not compute MRR, cluster ARI, NMI, V-measure, Kendall’s τ, or Spearman correlation for any bridge in this chapter. Keep this distinction sharp throughout: a generic metric being useful in principle is not evidence that this book measured it for this bridge.

Eight measurements, not a ladder

It is tempting to read the eight fields as increasingly strict rungs — reconstruction first, then neighborhoods, then task quality, then fine distinctions — as though a bridge that clears one automatically has a head start on the next. Resist that framing. These are distinct, non-nested views of the translated representation, not a monotonic sequence:

target-point reconstruction        — how close does the translated vector land to its target?
local neighborhood                 — do nearby items stay nearby?
relative ordering                  — do pairwise comparisons survive?
task retrieval                     — does end-to-end ranking quality survive?
calibration behavior               — does an absolute threshold still mean the same thing?
relation profile                   — do aggregate per-relation distances hold their shape?
hard-negative margin               — does one specific fine decision margin survive?
held-out generalization            — does performance on unseen entities match training performance?

A higher score on one does not mechanically imply a higher score on another, and the measured profile above already demonstrates disagreement among properties. But do not rank the eight measurements by raw numeric magnitude: 0.4658 reconstruction cosine, 0.9148 retrieval ratio, 0.8605 Pearson correlation, and 0.7566 calibration-transfer score are different statistics with different constructions and references. They also do not all use the same evaluation unit — item-level measurements use held-out items, retrieval uses qualifying held-out queries against the target test pool, and relation metrics use qualifying held-out pairs. Call this a measurement stack or a preservation profile, not a ladder or a leaderboard. Each entry is a separate sensor pointed at a different question, and none of them can vouch for what another one saw.

Defining each field from its implementation, not its name

Every field name below was chosen for readability, and every one of them invites a slightly more generous reading than its implementation supports. Define each precisely once, because an authorization decision — Chapter 20’s entire subject — is exactly the place where a mislabeled metric does real damage.

coordinate_reconstruction = 0.4658. Mean cosine similarity between each held-out item’s bridged vector and its corresponding native target vector, both L2-normalized. This is a generic reconstruction question, answered with a specific instrument (mean cosine); MSE is a different, legitimate reconstruction metric that this artifact does not compute.

neighborhood_at10 = 0.7434. Build a 10-NN graph among the bridged held-out vectors and a separate 10-NN graph among the native target held-out vectors; for each item, compute |N_bridge(i) ∩ N_native(i)| / 10; average over items. A mean top-10 intersection fraction, not Jaccard similarity (no union in the denominator) and not Recall@10 against relevance labels.

retrieval_ratio = 0.9148. Embed each qualifying held-out query in source space, translate it through the bridge, and score it against the native target test-item pool; compute nDCG@10; divide by the nDCG@10 a native target query achieves against the same pool. A task-quality ratio on one specific representation path — source query, translated, searched against a native target index. It is not top-k agreement, not Recall@10, not MRR, and not a measurement of the reverse path (a native target query against bridged source documents), which was never evaluated here.

rank_triplet_agreement = 0.7098. Sample triplets (a, b, c) from the held-out items; ask whether sim(a,b) > sim(a,c) holds with the same Boolean result in bridged space and in native target space; average agreement across sampled triplets. A sampled pairwise-order agreement rate, not Kendall’s τ and not Spearman rank correlation — both are legitimate alternative rank statistics that were not computed here.

calibration_transfer = 0.7566. Fit the approximate EER threshold on native-target held-out positive/negative pairs; apply that same numeric threshold, unchanged, to the bridged scores on the same pairs; compute the bridged FAR/FRR average; report max(0, 1 − (bridged_error − native_error)). This is an artifact-specific error-transfer score — not a probability, not a calibration percentage, and, because the formula clamps only from below, not bounded above by 1 either. The Wave 3 artifact contains a case where this score exceeds 1 (bge-large → mxbai-large, relative-representation rung: calibration_transfer = 1.026), because that bridge’s error came in below the native error at the native threshold. Read every value against its formula, not against an assumed [0, 1] range.

relation_profile_corr = 0.8605. Compute the mean cosine for each relation type (paraphrase, negation, entailment, and so on) separately in bridged space and in native target space, using only relations with at least three qualifying held-out pairs; take the ordinary Pearson correlation between the two resulting profiles. A correlation of two short vectors of means — not rank correlation, not “relation ordering,” and not relation-classification accuracy at the level of individual pairs. The next section shows exactly how much this can conceal.

hard_negative_ratio = 0.4735. For each qualifying held-out query, take the first grade-3 positive; among the query’s hard negatives, keep only those constructed by structured_perturbation; compute positive_score − max(negative_scores) — the margin — separately in native target space and in bridged space; average each; report bridged_mean_margin / native_mean_margin. This is a structured-perturbation mean-margin ratio, referenced against native TARGET, not source. It is not hard-negative agreement, not hard-negative accuracy, and — this distinction matters enough to repeat below — it is not measured against the source encoder’s own native performance at all.

ood_vs_id_reconstruction = 0.5562. The ratio of this same bridge’s coordinate-reconstruction score on held-out test entities to its coordinate-reconstruction score on its own training anchors. A held-out/train reconstruction ratio, not a reconstruction value in its own right — the artifact does not separately persist a training-set reconstruction figure, so nothing here supports a claim that training reconstruction was 1.00. The historical field name’s “OOD” means test entity families relative to training entity families under this corpus’s entity split — not open-world production distribution shift.

The first surprise: what retrieval_ratio does not say

Read retrieval_ratio = 0.9148 quickly, and it is easy to hear “91.5% of neighbors survived translation” or “91.5% of retrieval results agree.” Neither is what was measured. The number says: a bridged source-space query, searched against a clean native target item pool, achieved about 91.5% of the nDCG@10 a native target query achieves against that same pool, on held-out entity families. That is a real and useful fact about one specific data path. It does not measure or authorize whether bridged source documents would retrieve well under a native target query — the reverse path Chapter 20 already flagged as requiring its own separate evidence — and it does not measure whether the translated vector’s own top-10 neighbor set matches the native vector’s top-10 set, which is exactly what neighborhood_at10 measures instead, and exactly the distinction Chapter 22 builds an entire experiment around. The same underlying number, read under a looser name, would sound like a stronger and broader claim than the measurement supports.

The second surprise: what relation_profile_corr conceals

relation_profile_corr = 0.8605 is, on its face, reassuring — the aggregate shape of this bridge’s per-relation cosine profile tracks the native target’s shape fairly well. Look underneath it, using row 3.7’s per-relation detail for this same Procrustes bridge:

relation             native mean cosine    bridged mean cosine

paraphrase                  0.817                 0.726
negation                    0.784                 0.833

Two relations, each individually well within the range that produced the 0.8605 correlation. Now look at the contrast between them, which is the quantity a polarity-sensitive consumer would actually depend on:

paraphrase-minus-negation gap
  native:   0.817 − 0.784 = +0.033
  bridged:  0.726 − 0.833 = −0.107

Natively, paraphrase pairs sit slightly more similar than negation pairs — the direction a consumer distinguishing “this restates the claim” from “this negates the claim” would want. Under this bridge, that ordering reverses: negation pairs now average more cosine-similar than paraphrase pairs, on this held-out relation sample. The single correlation value 0.8605 does not identify which individual contrast changed sign; recovering that fact requires inspecting the relation-level components. A high Pearson correlation across an entire relation profile is a statement about the whole vector of means moving together — it is not a guarantee that any one specific contrast inside that profile survives, let alone survives its sign.

A high global correlation can coexist with a decision-relevant contrast changing sign.

This is the same shape of lesson the book has taught at several other layers: Chapter 9’s exact-versus-approximate distinction, Chapter 13’s aggregate metric versus its evaluation population, Chapter 14’s score versus its operating threshold, Chapter 15’s scalar score versus a richer signal bundle. Each time, an aggregate number that looked complete turned out to be silent about exactly the thing a specific consumer needed. Correlation is not indicted by this result — it correctly reports what it was built to report. The boundary condition is what needed stating, and now it has one.

Two disciplines belong beside this result, worth stating precisely rather than loosely:

Sign is not “more is better.” A positive bridged_minus_native shift for equivalent (+0.036) or temporal-mismatch (+0.168) is not automatically good news — it means those relation pairs became more cosine-similar under the bridge than natively, and whether that is desirable depends entirely on what a consumer expects that relation to look like. Averaging all nine relation shifts into one “how much improved” figure would erase exactly the distinction this section just surfaced.

Row 3.7 measures relation-level geometry, not classification. It reports mean-cosine shifts per relation, not per-instance relation-classification accuracy, not a confusion matrix, and not NLI correctness. The paraphrase-negation gap inversion above is a statement about aggregate cosine geometry for this held-out relation sample — a real and striking one — not a claim that every individual negation pair now outranks every individual paraphrase pair.

relation-swap was not measured in this slice. A natural further question — did A acquired B stay distinct from B acquired A? — is exactly the kind of thing this framework should be able to check. It cannot, here: the relation-preservation.json artifact only reports a relation type when at least three qualifying held-out pairs exist for it, and no relation-swap entry appears for this model pair under this split. The honest state is relation-swap: NOT MEASURED IN THIS SLICE, not a silently assumed pass or fail. Absence of evidence is not evidence of preservation, any more than it is evidence of failure.

One more piece of provenance worth naming precisely. Row 3.7 selects, per model pair, whichever row 3.4 rung had the highest neighborhood_at10 value, and reports that rung’s relation detail. For minilm-l6 → mpnet-base, that happens to be Procrustes — the same bridge this chapter has been auditing throughout, which is a fortunate coincidence for continuity, not a guarantee built into the artifact. The field is literally named best_rung, but “best” here means highest measured neighborhood overlap among the rungs row 3.4 evaluated — not best relation preservation, not best overall bridge, and not Chapter 20’s authorized choice. Read the selection criterion before trusting what “best” implies.

None of this licenses a verdict on the universal-geometry conjecture from Chapter 18. One synthetic corpus, one model pair, one method, one held-out split cannot prove or disprove a hypothesis about neural representations in general. What this evidence does support is narrower and still worth stating plainly:

The strongest version of a universal-geometry claim — one in which every relation-level distinction transfers uniformly — is not supported by this bridge. Any weaker version of the conjecture must be scoped to the specific properties actually shown to transfer, not assumed to cover every property a global correlation happens to summarize well.

The third surprise: a metric filename is a claim too

The historical artifact supervised-bridge-ceiling.json — Wave 3 row 3.6 — invites a specific reading before anyone opens the code: a supervised bridge, trained to translate all-MiniLM-L6-v2 vectors into all-mpnet-base-v2 space, tested against whether it can beat the untranslated source’s own hard-negative accuracy. That reading is wrong in a way that matters, and the only way to catch it is to read the implementation rather than trust the name.

What the code actually does. For each model pair, it collects (query, positive, structured-perturbation negative) triples from split_entity:train queries. It embeds the query with the source encoder. Both the positive and negative documents are drawn from p["Xs"] — source-space item vectors, not target-space vectors. It trains a small network, Linear(d_source, 512) → GELU → Linear(512, d_source), that maps a source-space query embedding to another vector of the same source dimensionality, L2-normalizes that output, and scores it by dot product against the same source-space positive and negative document vectors, optimizing a margin loss (ReLU(0.2 − (score_pos − score_neg))). At no point does the target encoder’s embedding matrix Xt participate in training or scoring. This is, precisely:

A supervised, nonlinear transformation of source-space query vectors, evaluated entirely within source space — not a cross-space bridge, and not a translation into target-space coordinates.

Call it a source-space readout control, not a “bridge.” The historical filename and the historical pair labeling ("minilm-l6 vs mpnet-base") name the target model only because that is which model pair the harness was iterating over when it ran this control for each source encoder — the target model’s embeddings are never consulted by the transformation itself.

What it measured, in full — not just the striking cell. Each source encoder gets its own 62 held-out structured-perturbation test triples:

source encoder     native source cosine accuracy    supervised source-readout accuracy

bge-large                    0.935                          0.935
minilm-l6                    0.952                          1.000
mpnet-base                   0.952                          0.919

Reporting only the MiniLM row would tell a cleaner but false story. The full table shows a supervised readout that helped one source encoder, left one exactly unchanged, and hurt one — on this one small held-out sample, in this one recipe. The defensible conclusion is an existence result, not a general recipe:

Native source-space cosine accuracy is not a universal upper bound on supervised readout performance: one fitted source-space query transform exceeded it on one held-out structured-negative test. Once its weights are fixed the transform is deterministic at inference, but this result is not a claim that the training procedure itself is deterministic or that the same improvement recurs across seeds. The same recipe did nothing for a second source encoder and made a third worse.

The real information boundary, stated precisely. A genuine limit does exist, and it is worth stating exactly rather than loosely: if two relevant inputs produce exactly the same representation, and a fixed deterministic downstream transformation receives only that representation, the transformation cannot map those two identical inputs to two different outputs. That is a hard constraint on the per-example information available to the readout. But native cosine performance is not the same quantity as information content. A supervised learner also receives task information during training, and that information is stored in its fitted parameters; what does not change at test time is the per-example source representation supplied to the learned readout. Row 3.6’s MiniLM result therefore supports the narrower claim: a learned readout can reorganize or expose distinctions that native cosine does not expose well, but it cannot distinguish two held-out inputs that are literally identical at its representation input unless some additional per-example signal is provided. It is not evidence about “the true information-theoretic ceiling” of any encoder, which this experiment does not attempt to measure and which is not directly measurable by cosine accuracy at all.

Do not compare row 3.6’s numbers with row 3.4’s hard_negative_ratio numerically. They are dimensionally different measurements, computed on different representations, against different references: row 3.4’s hard_negative_ratio is a ratio of margins, computed in bridged versus native TARGET space, using the Procrustes bridge under audit throughout this chapter. Row 3.6’s numbers are binary accuracies, computed entirely in source space, using an unrelated nonlinear query-readout network that this chapter’s Procrustes bridge never touches. A reader who subtracts or divides one against the other has manufactured a number neither experiment produced.

Artifact names are claims too. When implementation and filename disagree, the implementation wins. This is the same lesson Chapter 17 taught with calibrated_merge_ndcg10 (an offset that collapsed to zero under L2 normalization) and Chapter 19 taught with nonlinear-vs-linear-unpaired.json (an experiment run entirely on paired anchors). The recurrence is not an accident of this book’s evidence base — it is the discipline the book keeps teaching: read the code path that produced a number, not the label someone gave the file.

A separate diagnostic: the round trip

Row 3.9’s round-trip measurement was already introduced in Chapter 20 as a separate experiment from this chapter’s Procrustes bridge — independently fitted forward and reverse ridge maps, not the Procrustes pipeline under audit here. It is worth one more careful pass, because its two cosine columns are exactly the kind of similarly named, differently referenced quantity this chapter exists to catch:

pair                          forward-to-target cosine   return-to-source cosine   round-trip neighborhood@10

bge-large ↔ mxbai-large               0.8814                    0.8582                     0.6981
minilm-l6 ↔ mpnet-base                 0.5860                    0.7940                     0.7678
mpnet-base ↔ bge-large                 0.7893                    0.7449                     0.7314

The forward column compares the forward ridge map’s output against the native target vector. The round-trip column compares the round-tripped result — translate to target, then translate back — against the native source vector it started from. These are different reference objects, evaluated against two different maps composed together, and for minilm-l6 ↔ mpnet-base the round-trip figure is numerically higher than the forward one — not a paradox, only two incomparable quantities placed in adjacent columns.

Two similarly named numbers can be incomparable if their reference objects differ. Subtracting single_hop_forward_cosine from roundtrip_cosine, or the reverse, produces a difference with no defined meaning; do not compute it, and do not read this table as showing “loss per hop.” The safe, narrow lesson stands exactly where Chapter 20 left it: a separately fitted round trip does not perfectly reconstruct the source, and a round trip is its own transformation path requiring its own end-to-end evaluation — never inferred from either leg’s own score in isolation.

Every metric needs its own reference — there is no universal baseline column

Several of this chapter’s numbers are constructed relative to a native target figure, and it is tempting to read every metric as “distance from 1.0” or “how close to perfect.” Neither reading is safe. retrieval_ratio can exceed 1 when a bridge happens to outperform native on the measured task — the artifact contains exactly this for bge-large → mxbai-large under Procrustes (retrieval_ratio = 1.0035). calibration_transfer can exceed 1 for the same structural reason, as the relative-representation example above showed. Native target is a comparison reference for the specific property it measures against, not a mathematical ceiling every metric approaches from below.

Different reference controls answer different questions, and none of them substitutes for another:

Native target — what the target representation does natively, on the same held-out population. The comparison point most of row 3.4’s ratios are built against.

Null-map (direct coordinate reuse) control — measured only where mechanically defined (matching widths), as in Chapter 19’s bge-large → mxbai-large control (reconstruction 0.9809, neighborhood 0.8818, retrieval ratio 0.9952). A strong null result establishes that direct reuse preserved the measured property well on that pair and workload — it does not establish shared training regime, intentional pre-alignment, or a common space_hash. Two distinct declared identities remain distinct regardless of how well direct reuse happens to score.

Random-map control — a genuinely useful sanity check (does the bridge do better than a random transformation of the same shape?) and a recommended, proposed control, not one present in the current Wave 3 artifacts. Do not report a random-map number in this chapter’s tables; none exists.

Round-trip diagnostic — as above, a separate composed-path measurement with its own reference frame, not a floor and not a bound on any other bridge’s composition behavior.

Source-native readout control — row 3.6, a source-space-only experiment using a different transformation entirely, informative about whether native cosine is a ceiling for a different question than any row 3.4 field measures, and never a baseline for hard_negative_ratio.

A reference must match the quantity being measured. There is no single baseline column that makes sense for every preservation metric at once. Reach for the reference each metric’s own construction demands, not a generic “how far from perfect” intuition, and not a raw cross-model cosine — Chapter 16 already established that a raw cross-model cosine between two unaligned spaces carries no authorized semantic interpretation in either direction, so it cannot serve as a floor for anything measured in this chapter.

Reading a pattern without issuing a verdict

Certain shapes recur often enough across this book’s evidence to be worth naming as things to investigate next, deliberately short of a pass/fail rule:

PatternWhat the measurements suggestWhat to test next
high reconstruction, weak neighborhoodthe point lands near its target, but local relational structure differsa neighborhood-sensitive consumer’s actual requirement
high neighborhood, weak rank-triplet agreementthe local set survives better than pairwise order within ita ranked-list consumer’s actual requirement
good retrieval ratio, weak calibration transfertask ordering may survive while absolute score behavior changesa thresholded consumer’s own calibration contract (Ch14)
high relation-profile correlation, one contrast reversedthe global profile can hide a relation-specific failurethe exact relation contrast the consumer depends on

This table names hypotheses to test, not verdicts to hand down — Chapter 20 owns the machinery for turning a named requirement plus a cited observation into ALLOW, CONDITIONAL, DENY, or NOT_EVALUATED. A consumer’s actual requirement may need one metric, or several at once, or a specific slice of one metric (a particular relation, a particular query population); nothing in this chapter’s evidence forces every consumer into a single deciding column. Define the requirement first — precisely, and before looking at the results — then inspect exactly the measurements that requirement names.

What this chapter establishes and what it does not

Establishes: what each of Wave 3’s eight scalar preservation summaries actually computes, sourced from its implementation rather than its name, while preserving the per-relation detail row 3.4 also records; that these properties form a distinct, non-nested measurement stack, not a strict ladder; that retrieval_ratio describes one specific representation path and does not measure neighborhood identity or the reverse retrieval direction; that a strong global relation_profile_corr can coexist with a specific relation contrast (paraphrase-minus-negation, +0.033 native → −0.107 bridged) reversing sign; that relation-swap was not measured in this held-out slice, and absence of a measurement must be reported as such rather than assumed either way; that the artifact historically named supervised-bridge-ceiling.json is a source-space-only supervised query-readout experiment, not a cross-space bridge, and that it beat native source cosine on one of three source encoders while leaving one unchanged and worsening a third; that native cosine performance and information content are different quantities, with a precise and narrower boundary case supporting that claim; that the round-trip diagnostic’s two cosine columns reference different endpoints and cannot be subtracted; and that every metric needs a reference control matched to what it actually measures — native target, a null-map control where mechanically defined, a round-trip diagnostic, or a source-readout control — with no single universal baseline serving every metric.

Does not establish: a single “preservation score” for any bridge (there is not one); that this or any bridge is authorized for any specific operation (that is Chapter 20’s separate machinery, requiring a stated consumer requirement this chapter does not supply); that the universal-geometry conjecture is proven, disproven, or scoped specifically to “coarse” relations as a general rule (one bridge’s relation profile does not settle a general hypothesis); that supervision generally helps, generally hurts, or generally does nothing to a source-space readout (the evidence here shows all three outcomes across three source encoders); statistical significance or uncertainty for any reported value (Wave 3 rows 3.4, 3.6, 3.7, and 3.9 report point estimates only, with no bootstrap interval, repeated split, or significance test); or that this chapter has separated counterpart recovery from structural fidelity in the sense Chapter 22 measures directly — it has not, and that gap is exactly where this chapter’s evidence stops.

Lab 21: audit the evidence, one field at a time

MEASURED — artifacts ladder-8property-matrix.json (row 3.4), relation-preservation.json (row 3.7), supervised-bridge-ceiling.json (row 3.6), roundtrip.json (row 3.9). REPRODUCIBLE — python run_wave3.py 3.4 3.6 3.7 3.9.

Question. For the MiniLM→mpnet Procrustes bridge, what does each preservation field actually know — and what is it structurally unable to see?

Step 1 — load the row 3.4 evidence for this exact bridge, and define every field from its implementation before reading anything into its value:

coordinate_reconstruction              0.4658
neighborhood_at10                      0.7434
retrieval_ratio                        0.9148
rank_triplet_agreement                 0.7098
calibration_transfer                   0.7566
relation_profile_corr                  0.8605
hard_negative_ratio                    0.4735
ood_vs_id_reconstruction               0.5562

Step 2 — ask what looks reassuring, and resist concluding anything yet. retrieval_ratio = 0.9148 and relation_profile_corr = 0.8605 both look strong. Record that observation without converting it into a pass — no consumer requirement has been named, and Chapter 20 already established why that matters.

Step 3 — attack the global relation metric with row 3.7’s detail. For this same bridge (Procrustes was selected as the best_rung by highest neighborhood_at10):

native paraphrase mean cosine    0.817        bridged paraphrase mean cosine    0.726
native negation mean cosine      0.784        bridged negation mean cosine      0.833

native paraphrase-negation gap  +0.033        bridged paraphrase-negation gap  -0.107

State the finding exactly: the global relation-profile Pearson correlation did not expose this contrast reversal.

Step 4 — inspect every relation mean shift, and resist scoring positive values as improvements:

contradiction       -0.020
entailment          -0.026
entity-related       -0.016
equivalent           +0.036
negation             +0.049
paraphrase           -0.091
partial-support      -0.030
temporal-mismatch    +0.168
topic-related        +0.007

Step 5 — identify missing evidence rather than inventing it. relation-swap does not appear among the relations reported for this pair in relation-preservation.json (fewer than three qualifying held-out pairs). Record it as NOT MEASURED IN THIS SLICE, not as a silent pass or fail.

Step 6 — audit the hard-negative field. hard_negative_ratio = 0.4735 is the ratio of bridged to native-target mean structured-perturbation margin — not accuracy, not agreement, and not comparable to any row 3.6 figure.

Step 7 — audit calibration transfer. calibration_transfer = 0.7566 is the artifact-specific error-transfer score defined above. State its formula; issue no pass or fail — no calibration requirement has been declared for this bridge’s consumers.

Step 8 — run the separate source-cosine readout control. From row 3.6, all three source encoders:

bge-large         0.935 -> 0.935
minilm-l6         0.952 -> 1.000
mpnet-base        0.952 -> 0.919

62 held-out structured-perturbation triples per source. State plainly that this is a source-space-only experiment, not this chapter’s cross-space bridge.

Step 9 — audit the round trip separately. From row 3.9, confirm that single_hop_forward_cosine (reference: native target) and roundtrip_cosine (reference: native source) use different endpoints and must not be subtracted or otherwise combined.

Step 10 — write the conclusion as metric boundaries, not scores:

retrieval_ratio:
  measures:      relative task nDCG@10 on one representation path (source query -> bridge -> native target pool)
  does not measure: neighborhood identity, counterpart recovery, or the reverse retrieval direction

relation_profile_corr:
  measures:      aggregate similarity of two per-relation mean-cosine profiles (Pearson)
  does not measure: preservation of any single named contrast within that profile

hard_negative_ratio:
  measures:      bridged/native-target mean structured-perturbation margin ratio
  does not measure: source-native accuracy, or hard-negative agreement in any classification sense

calibration_transfer:
  measures:      an artifact-specific error-transfer score at the native target's EER threshold
  does not measure: a percentage of "threshold preserved," and is not bounded above by 1

roundtrip_cosine:
  measures:      return-to-native-source cosine after two independently fitted composed maps
  does not measure: a "two-hop loss" relative to the single-hop forward cosine, which references a different endpoint

Try it yourself

Take one bridge you have evidence for. For every field in its preservation profile, write its exact computation — not its name — before writing anything about what the number means. Then pick one relation, contrast, or slice your own downstream consumer actually depends on, and check whether an aggregate metric in the profile could plausibly be hiding a reversal in exactly that slice, the way relation_profile_corr did here. If a generic question you care about (counterpart recall, cluster agreement, rank correlation) has no corresponding measured field, record it as NOT MEASURED, not as a pass or fail by default.

Companion component: the preservation observation

The output of this chapter’s audit is evidence, precisely defined and bound to its own reference — never a verdict. It stays strictly upstream of Chapter 20’s authorization layer:

preservation_observation:
  bridge_id:
  bridge_version:

  identities:
    source_space_hash:
    target_space_hash:
    derived_output_space_hash:
    direction:

    evaluation_contract_ref:
  corpus_hash:
  split_ref:                     split_entity: train / test
  uncertainty_status:            point_estimates_only   # no bootstrap CI / repeated split / significance test

  measurements:
    coordinate_reconstruction:
      value:              0.4658
      definition:          mean cosine, bridged item vs native target item
      population:          split_entity:test items
      reference:            native_target

    neighborhood_at10:
      value:              0.7434
      definition:          mean 10-NN intersection fraction, held-out items
      population:          split_entity:test items
      reference:            native_target

    retrieval_ratio:
      value:              0.9148
      definition:          bridged-query nDCG@10 / native-target-query nDCG@10
      population:          qualifying held-out queries; candidate pool = split_entity:test items
      representation_path:  source_query -> bridge -> native_target_item_pool
      reference:            native_target

    rank_triplet_agreement:
      value:              0.7098
      definition:          sampled Boolean triplet-order agreement
      population:          triplets drawn from split_entity:test items
      sampling_method:      deterministic RNG draws (seeds 0/1/2), up to 300 draws, non-distinct triplets discarded

    calibration_transfer:
      value:              0.7566
      definition:          max(0, 1 - (bridged_error - native_error)) at native EER threshold
      population:          held-out relation pairs used by the EER contract
      native_threshold_ref: row_3.4_native_eer
      bounded_above:        false

    relation_profile_corr:
      value:              0.8605
      statistic:            pearson
      population:          relation types with >=3 qualifying split_entity:test pairs
      reference:            native_target_per_relation_means

    hard_negative_ratio:
      value:              0.4735
      negative_method:      structured_perturbation
      population:          qualifying held-out queries with first grade-3 positive in test and available structured perturbations
      reference_positive_policy: first_grade3
      reference:            native_target_mean_margin

    heldout_train_reconstruction_ratio:
      value:              0.5562
      definition:          test coordinate_reconstruction / train coordinate_reconstruction
      populations:         split_entity:test items / split_entity:train anchors

  relation_detail:
    selected_rung:         procrustes
    selected_rung_reason:   highest neighborhood_at10 among row 3.4 rungs for this pair
    native_means:           { paraphrase: 0.817, negation: 0.784, ... }
    bridged_means:          { paraphrase: 0.726, negation: 0.833, ... }
    bridged_minus_native:   { paraphrase: -0.091, negation: +0.049, ... }
    specific_contrasts:
      paraphrase_minus_negation: { native: +0.033, bridged: -0.107 }
    unmeasured_relations:   [ relation-swap ]

  auxiliary_controls:
    null_map_ref:                       not_applicable   # dims differ, 384 vs 768
    roundtrip_ref:                      row_3.9 (independent ridge maps, not this bridge)
    source_cosine_readout_control_ref:  row_3.6 (source-space only, not this bridge)
    random_map_ref:                     not_measured   # recommended, not present in Wave 3

  provenance:
    artifact_refs:          [ladder-8property-matrix.json, relation-preservation.json]
    code_hash:

Notice what this record deliberately excludes: no verdict, no usable_for, no consumer_metric, no PASS/FAIL. Those fields belong to Chapter 20’s authorization layer, which cites this observation but is never produced by it. Chapter 21 produces evidence. Chapter 20 decides how evidence becomes authorization. Keeping the two artifacts structurally separate is not bureaucratic caution — it is the only way to prevent exactly the failure this chapter spent its middle sections diagnosing: a number that looked complete standing in for a decision it was never built to make.

Embedding Observatory progression

By the end of this chapter, the Observatory should be able to answer, for any preservation number a reader or a system encounters: what exactly was computed — a raw score, a ratio, a correlation, or an error-derived score? Which two sets of vectors were compared, and in which representation? What was the reference — native source, native target, a training split, or another derived representation entirely? Which population produced it — all held-out items, qualifying queries only, or a specific relation subset? Which representation path did it evaluate, and is that the path the current consumer actually needs? Which relations were present in a relation-level measurement, and which were silently absent? Which rung or method actually produced this observation, and by what selection criterion? Does the artifact’s filename or field name match its implementation, or does the code say something different? What failure mode is this specific measurement structurally incapable of detecting? And is any uncertainty estimate available, or is this a single point estimate on one frozen split?

This is the natural companion to Chapter 20’s question. Chapter 20 asks why is this operation allowed. Chapter 21 asks what exactly does the evidence behind that answer mean — and, just as importantly, what it cannot mean, no matter how the field happens to be named.

Failure modes

  • Calling a measurement a verdict. Chapter 20 separated evidence from authorization; this chapter’s eight fields are evidence, not eight verdicts.
  • Trusting the friendly metric name. retrieval_ratio is not retrieval agreement; hard_negative_ratio is not hard-negative accuracy; relation_profile_corr is not relation ordering.
  • Treating a strong relation-profile correlation as protection against a specific contrast failing. 0.8605 coexisted with the paraphrase-negation gap flipping from +0.033 to -0.107.
  • Treating a signed relation-cosine shift as automatically an improvement. Sign means something different for paraphrase than for negation or temporal-mismatch.
  • Comparing hard_negative_ratio numerically against row 3.6’s source-native accuracy. Different measurements, different representations, different references — never subtract or divide one by the other.
  • Calling row 3.6 a MiniLM→mpnet bridge. It transforms and scores entirely within source space; the target encoder’s vectors never participate.
  • Trusting supervised-bridge-ceiling.json’s filename literally. The historical name does not override what the code actually computes.
  • Treating native cosine accuracy as a measurement of information content. It is one specific readout; a different, supervised readout can sometimes expose a distinction native cosine does not.
  • Assuming supervision always helps. It helped one source encoder, left one unchanged, and hurt one, in this one recipe on this one held-out sample.
  • Using a raw cross-model cosine as a baseline. Chapter 16 already ruled this out — such a value has no authorized semantic interpretation on its own.
  • Calling a null-map result a floor. Where mechanically defined, a strong null-map score is a strong control, not evidence of a lower bound or a claim that two distinct space identities collapsed into one.
  • Reporting a random-map result that was not measured. It remains a recommended, proposed control absent from the current Wave 3 artifacts.
  • Calling round-trip cosine a “two-hop loss” relative to forward reconstruction cosine. The two columns reference different endpoints and cannot be subtracted.
  • Issuing PASS/FAIL on a preservation observation with no stated consumer requirement. That decision belongs to Chapter 20’s authorization layer, not to this chapter’s evidence record.
  • Reporting relation-swap preservation when it was not measured for this slice. Missing evidence must be reported as missing, not silently assumed either way.
  • Generalizing calibration transfer as universally poor. It varies substantially by pair and by method, and this artifact even contains a value above 1.

What this chapter established

  • Chapter 20 made preservation evidence operationally load-bearing; this chapter audited what that evidence actually says, field by field, from each metric’s implementation rather than its name. A preservation metric measures one property, on one representation path, against one specific reference — never bridge quality in general.
  • Row 3.4’s eight scalar summaries form a distinct, non-nested measurement stack, not a ladder. Each field’s exact definition differs from what its friendly name suggests, and generic alternatives (MSE, Recall@k, MRR, Kendall’s τ, cluster ARI, paired-target counterpart recovery) are useful in principle but were not measured here.
  • A strong global correlation can conceal a decision-relevant contrast flipping sign. relation_profile_corr = 0.8605 coexists with the paraphrase-minus-negation gap reversing from +0.033 native to −0.107 bridged. Relation-swap was not measured in that held-out slice, and the absence was recorded rather than papered over.
  • supervised-bridge-ceiling.json is a source-space-only query-readout experiment that never touches the target encoder’s vectors. It beat native source cosine for one of three encoders, left one unchanged, and worsened one: native cosine is not a universal upper bound on supervised readout, and supervision does not automatically help.
  • No single reference control serves every metric — native target, a mechanically-defined null map, a round-trip diagnostic, and a source-readout control each answer a different question, and the round trip’s two cosine columns reference different endpoints, so they must never be subtracted as a “loss per hop.” The preservation observation records definitions, references, populations, and explicitly unmeasured properties, deliberately excluding any verdict.

Next

Chapter 21 audited eight measurements and found that even a well-defined field can be misread if its reference and representation path go unexamined. But one deceptively simple question never got asked directly by any of Wave 3’s eight fields: if a bridge translates one object, can it retrieve that object’s own native target counterpart — not just achieve good aggregate task nDCG, and not just preserve neighborhood overlap in the abstract, but land close enough to its specific paired target to recover it? That is counterpart recovery. It is not the same question as structural fidelity — did the translated point reproduce the geometry that actually surrounds its native target, correct counterpart or not? The preservation profile taught this chapter to separate properties. Chapter 22 reveals that even the word “retrieval” still hides two different contracts inside it — recovering the right point, and recovering the shape of the space around it — and measures both directly, on a benchmark built to tell them apart.