← Embeddings From First Principles

Retrieval Is Not Geometry

A bridge can recover the exact paired target almost every time while reproducing markedly less of the neighborhood around it. Separate counterpart recovery from structural fidelity on the same translated vectors, attack the interpretation with a high-CKA control and a shuffled-correspondence control, and learn that even a well-defined metric can fail to discriminate what another one discriminates clearly.

Part VI — Crossing Embedding Spaces · Did it find the counterpart, or recreate the neighborhood?

The result that looks like success

Fit a bridge, translate a held-out source vector, and ask the simplest possible question: does the exact native target counterpart show up nearby? On the Wave 6 benchmark, for a paired ridge bridge translating mxbai-embed-large-v1 into Qwen3-Embedding-8B, the answer is about as good as that question can get. Across 589 held-out sentences, using the same fitted bridge and the same translated vectors throughout:

paired target in top 10                    1.0000    (Recall@10)
paired target is the top-ranked result     0.8676    (Recall@1)
mean reciprocal rank of the paired target   0.9221    (MRR)

That looks finished. Now ask a different question of the same translated vectors, deliberately excluding the paired target itself: does the translated point’s neighborhood — the other nearby target items — resemble the native neighborhood around the paired target?

top-10 neighborhood overlap (counterpart excluded)   0.7964
local shared-neighbor rank correlation (Spearman)     0.6987
cluster-assignment agreement (ARI)                    0.4750

Top-10 counterpart recovery is perfect on this held-out pool, and top-1 recovery is strong but not perfect. The structure around the paired target is reproduced less faithfully under the neighborhood and local-order measurements. This is not a failure — it is a demonstration that the first question and the second question are not the same question, even though they are computed from the exact same translated output.

Recovering the right point is not the same thing as preserving the relationships around that point.

Chapter 21 taught that a metric’s friendly name can hide what it actually computes. This chapter applies the same discipline one level up: the word “retrieval” itself hides at least two different contracts, and a bridge can satisfy one almost perfectly while only partly satisfying the other.

Two contracts hidden inside one word

Counterpart recovery. For an object x, translate its source vector through the bridge and ask whether the native target vector for that same object — not a similar object, the exact paired one — is what the translated point finds:

x
│
├─ E_A(x) ── T ──> translated point in B
│
└─ E_B(x) ───────> native target point in B, the object's OWN counterpart

Measured by: cosine_to_target, recall_at_1, recall_at_10, mrr — all defined against one specific reference vector per object, E_B(x).

Structural fidelity. Set the paired counterpart aside entirely and ask about the translated point’s relationship to the other target-space items around it:

N_native(x)      = native target-space neighbors of E_B(x), excluding E_B(x) itself
N_translated(x)  = target-space neighbors around T(E_A(x)), with E_B(x) explicitly excluded

Measured by: agreement_at_10, order_preservation_local (a Spearman correlation), triplet_agreement, and cluster_preservation_ari. The first three probe relationships among the translated point, the native target point, and other target items; cluster_preservation_ari instead compares cluster assignments over the translated and native test matrices as wholes.

These are genuinely different questions. Counterpart recovery constrains one relationship per object — T(A_i) landing near B_i. Structural fidelity constrains many relationships at once — how T(A_i) compares to nearby B_j relative to how B_i itself compares to those same target items. A fitted map can satisfy the first requirement much more strongly than the second: learning to place a held-out source vector near its paired-target location does not, by itself, force all of the surrounding target relationships to match.

A tiny example, with the counterpart properly separated

The earlier illustration risked collapsing the two contracts by mixing the paired object into the neighbor list being scored. Keep them apart, matching exactly how Wave 6’s agreement_at_10 is actually computed — with the counterpart excluded before the overlap is measured:

paired target:  A

native neighbors around A (A itself excluded):
  1. B
  2. C
  3. D
  4. E
  5. F

translated query ranking (unfiltered):
  1. A        <- the paired counterpart, recovered at rank 1
  2. B
  3. D
  4. E
  5. G
  6. H
  ...

Recall@1 for this object is a perfect success — A is the top-ranked result. Now, following Wave 6’s own construction, remove the paired counterpart A from the translated ranking before comparing neighborhoods:

translated neighborhood (A excluded):   B, D, E, G, H
native neighborhood:                     B, C, D, E, F

overlap:  {B, D, E} ∩ {B, C, D, E, F}  =  3 shared items  →  3 / 5

Counterpart recovery: perfect. Neighborhood overlap on the same object: three of five. Nothing about A sitting at rank 1 gave the neighborhood-overlap computation any credit — the exclusion rule means a perfect counterpart recovery and a mediocre neighborhood score can coexist on the very same translated point, exactly as the measured Wave 6 profile shows at scale.

The implication this experiment actually rejects

It is tempting to assume that a bridge which reproduces a target’s neighborhood well must also recover the paired counterpart well — after all, if T(A_i) sits among B_i’s true neighbors, surely B_i itself is nearby too. That intuition is not something these metrics guarantee, and this chapter’s evidence does not test it either way: Wave 6 measures pairs where recovery is strong, not a case engineered to test whether strong structural fidelity without strong recovery is possible. What the measured evidence does establish is the direction that matters for this chapter, stated as a counterexample to the naive assumption rather than a general theorem:

Excellent counterpart recovery does not imply excellent structural fidelity. Wave 6 supplies the concrete case: Recall@10 of 1.0000 coexists with 10-NN agreement of only 0.7964 and local order correlation of only 0.6987, on the identical translated vectors.

Neither contract logically contains the other. Do not read this chapter as establishing that good neighborhood preservation “usually” implies good counterpart recovery, or as establishing any general implication in that direction — no such claim is supported by what was measured.

Demonstration: the Wave 6 cross-space benchmark

MEASURED — artifact experiments/embeddings-from-first-principles/wave6/artifacts/cross-space-benchmark.json.

Corpus and split. 2,925 distinct short sentences (content_sha256 = 0c47690b068f7b07328fabea1fb9aa05af6e067d5b8160792dad588f71b35517), drawn from RELATE v0.1/v0.2 and RELATE-DOC v0.1 sentence material — a designed, templated corpus built for this book’s experiments, not a sample of naturally occurring text. Split deterministically by item ID: sha1(id) % 5 == 0 → test, giving 2,336 training sentences and 589 held-out test sentences. This is an item-ID split, not an entity-family split — a different partitioning rule from split_entity used throughout Chapters 18–21, and worth keeping distinct in memory when comparing evidence across chapters.

Encoders. Source: mixedbread-ai/mxbai-embed-large-v1, 1024-d, a sentence-transformers encoder. Target: Qwen/Qwen3-Embedding-8B, 4096-d, a decoder-LM embedding model — a deliberately different training regime from the source, chosen to make the cross-space question genuinely hard rather than to isolate any specific architectural cause. Control: BAAI/bge-large-en-v1.5, 1024-d. Every cached embedding matrix is L2-normalized before fitting and scoring. This is one intentionally cross-regime pair plus one high-structural-agreement control — not evidence about encoder families in general, arbitrary cross-modal alignment, or production-scale language coverage.

The headline map. Ridge(alpha=1.0), scikit-learn’s default fitted intercept, trained on L2-normalized source and target training vectors, predictions L2-normalized before scoring:

T_raw(x) = xW + b
T(x)     = normalize(T_raw(x))

This is the same affine ridge map family already evaluated in the preceding bridge chapters — familiar machinery, but now applied to a larger Wave 6 benchmark and a new source/target pair. No new map family is introduced here.

Headline result — mxbai → Qwen, n_train = 2,336:

counterpart recovery
  cosine_to_target           0.8837
  recall@1                   0.8676
  recall@10                  1.0000
  mrr                        0.9221

structural fidelity (same translated vectors, counterpart excluded from neighborhoods)
  agreement@10               0.7964
  local order (Spearman)     0.6987
  triplet agreement          0.9197
  cluster ARI                0.4750

Read the counterpart-recovery block carefully before moving on. recall_at_10 = 1.0000 means: for every one of the 589 held-out sentences, the exact paired native target vector appears somewhere in the top ten when the translated source vector ranks the full native target test pool by cosine. That is an excellent and precise fact. It does not, by itself, establish a coherent geometric region, preserved local topology, a preserved semantic cluster, or anything about production retrieval behavior — “the bridge always reaches the correct region” is an intuition worth holding only after stating the exact top-10-membership fact it stands in for, not a substitute for it. recall_at_1 = 0.8676 means the paired target is the single highest-scoring candidate for 86.76% of held-out objects — not 86.8% semantic correctness, and not a claim about whether any given object’s neighborhood is right. mrr = 0.9221 is the mean reciprocal rank of the paired target across all 589 objects, a natural companion statistic to Recall@1 and Recall@10 that describes how far down the ranking the misses tend to fall — it belongs squarely to counterpart recovery, not to any geometric claim.

Now read the structural-fidelity block with equal precision. agreement_at_10 = 0.7964 is the mean top-10 neighborhood intersection fraction, computed with the paired counterpart explicitly excluded from both neighbor lists before the overlap is measured — roughly, on average, four out of every five native neighbors survive translation, and the fifth is replaced by something else. order_preservation_local = 0.6987 is a mean Spearman rank correlation, but the implementation is more specific than the shorthand “shared-neighbor order” suggests. For each item it forms a local candidate set consisting of the native top-10 neighbors plus any translated top-10 items that also fall inside the native top-20 window, then compares the native-target cosine ranking and translated-point cosine ranking over that same candidate set. It is therefore not a correlation over only the intersection of the two top-10 lists, and it is not “70% of the ordering survives” or “70% order accuracy”; it is a rank-correlation coefficient describing how similarly those selected local candidates are ordered. triplet_agreement = 0.9197 is the fraction of sampled (a, b, c) item triples for which the cosine preference — is a closer to b or to c? — has the same sign natively and after translation; it is a broader, globally sampled pairwise-order statistic, distinct from the local Spearman statistic above, and the two need not move together. cluster_preservation_ari = 0.4750 is the Adjusted Rand Index between a KMeans(k=25, n_init=10, random_state=0) clustering of the native target test matrix and the same clustering algorithm applied to the translated test matrix — ARI is not a percentage-preserved figure; it is chance-adjusted agreement between two cluster assignments, where 1.0 means identical assignment up to label permutation, values near 0 indicate agreement no better than chance, and negative values are possible.

Do not build a monotonic staircase out of these five numbers — 0.475 < 0.699 < 0.796 < 0.868 < 1.000 — as though larger or more relational structure survives progressively less. Recall, a correlation coefficient, an intersection fraction, and an adjusted agreement index are different statistics on different scales, computed against different reference constructions; their raw magnitudes are not comparable positions on one shared fidelity axis. The qualitative result stands on its own without that false ordering: counterpart-recovery metrics are extremely strong on this pair; the two local-neighborhood measurements (agreement, local order) show materially less fidelity; cluster ARI is its own distinct statistic and must be read on its own terms, which the next section takes up directly.

Retrieval quality is itself plural

Chapter 21 established that Wave 3’s retrieval_ratio measured a bridged-source-query nDCG@10 divided by a native-target-query nDCG@10 — a task-quality ratio — and explicitly did not ask whether a translated item retrieves its own paired target. This chapter’s counterpart-recovery metrics fill exactly that gap. Laid side by side, “retrieval” now names at least three distinct contracts this book has separately measured:

1. counterpart recovery (this chapter, Wave 6)
   Does T(E_A(x)) recover its own paired E_B(x)?
   -> recall@1, recall@10, mrr, cosine_to_target

2. structural / neighborhood fidelity (this chapter, Wave 6)
   Does T(E_A(x)) reproduce the relationships surrounding E_B(x)?
   -> agreement@10, local order, triplet agreement, cluster ARI

3. task retrieval quality (Chapter 21, Wave 3)
   Does a full ranked retrieval workload preserve relative nDCG?
   -> retrieval_ratio, against a relevance-labeled query set

“Retrieval works” is not a complete sentence. Retrieved what, relative to which reference, for which consumer?

A migration workflow whose entire declared requirement is locating each object’s corresponding record needs only the first contract. A RAG context-selection service that depends on the set of retrieved evidence resembling what the native target model would have surfaced needs the second. A workload evaluated against relevance judgments across a full query population needs the third. None of the three substitutes for either of the others, and a single number — however good — never answers all three questions at once.

Attacking the interpretation with a high-structural-agreement control

A useful positive control asks whether the structural measurements rise on a pair whose native representations are already known to agree much more strongly. Chapter 16 found very high linear CKA for mxbai-embed-large-v1 versus bge-large-en-v1.5 on its earlier corpus; Wave 6 independently measures native test-pool CKA 0.9844 for that pair, against 0.8707 for mxbai versus Qwen. CKA is deliberately insensitive to orthogonal changes of basis, so this is evidence of strong structural agreement, not evidence of a shared space_hash or coordinate identity. Fit the same ridge construction to the high-CKA pair — a genuinely fitted bridge between two distinct encoder identities, not a null map or a raw-coordinate-reuse control — and evaluate it exactly the same way:

mxbai -> bge-large, paired ridge (n_train = 2,336), native CKA 0.9844

counterpart recovery
  cosine_to_target           0.9727
  recall@1                   0.9932
  recall@10                  1.0000
  mrr                        0.9966

structural fidelity
  agreement@10               0.9127
  local order (Spearman)     0.8944
  triplet agreement          0.9722
  cluster ARI                0.4761

Compare directly against the headline pair:

                    mxbai -> Qwen     mxbai -> bge (control)     difference

agreement@10           0.7964             0.9127                 0.1163
local order             0.6987             0.8944                 0.1957
cluster ARI             0.4750             0.4761                 0.0011

Two of the three structural-fidelity measurements discriminate this control sharply from the headline pair — agreement@10 and local order are both substantially higher on the structurally closer pair, consistent with the intuition that a more structurally similar target space should be easier to reproduce faithfully once a bridge lands near the correct point. The third — cluster_preservation_ari — barely moves at all: 0.4750 against 0.4761 is, for practical purposes, the same point estimate.

This is exactly the kind of result Chapter 21 taught this book to take seriously rather than explain away:

A positive control can audit the sensor as well as the system. If a metric fails to distinguish two cases that other well-defined structural measurements separate cleanly, that metric deserves less interpretive weight in this particular comparison — not a retroactive story about why it must be noisy.

Do not claim the control “preserves all three” structural measurements — it does not; cluster ARI does not separate the two cases at all. And do not diagnose cluster ARI itself as “noisy” as an established fact: this run used one fixed KMeans configuration (k=25, n_init=10, random_state=0) on one frozen train/test split, with no repeated seeds, no k sweep, and no bootstrap resampling to actually measure variance. The honest statement is narrower and still useful:

In this frozen run, cluster ARI does not distinguish the headline mxbai → Qwen bridge from the high-structural-agreement mxbai → bge control, even though agreement@10 and local order both do. That makes ARI weak evidence for this particular comparison, and a natural candidate for a sensitivity study — varying k, the KMeans seed, or the test partition — that this chapter’s evidence does not include.

Nor does the control license a claim about the general relationship between native structural similarity and bridge fidelity. Two model pairs, at two CKA values, is two observations — not a curve, not a scaling law, and not evidence that native CKA causes the fidelity differences observed. The two measured pairs are consistent with higher native structural agreement accompanying stronger neighborhood and local-order preservation; nothing here isolates causation, and no functional relationship between CKA and fidelity should be inferred from two points.

A negative control: destroying the correspondence

If the previous control shows what a structurally favorable pair looks like, a second control asks a sharper question: how much of the headline result depends on the training pairs actually corresponding to the same objects at all? Wave 6 fits the identical ridge construction to mxbai → Qwen, but with the training target rows randomly permuted under a fixed seed — the test-time correspondence remains correct for evaluation, only the training supervision is scrambled:

mxbai -> qwen, paired ridge, SHUFFLED training correspondence

cosine_to_target            0.4813
recall@1                    0.0034
recall@10                   0.0102
mrr                         0.0111
agreement@10                0.0421
local order (Spearman)     -0.1879
triplet agreement            0.6485
cluster ARI                  0.0729

Counterpart recovery collapses almost to nothing — Recall@1 of 0.34%, Recall@10 of 1.02%, both near what a random ranking over the test pool would produce. Neighborhood agreement falls to near zero, and local order goes slightly negative — a mean Spearman correlation below zero in this run, worth reporting exactly as that rather than reading any further meaning into it. Triplet agreement, interestingly, stays well above chance (0.6485) — a broad, globally sampled ranking statistic apparently retains some signal even under scrambled training correspondence, in a way the finer local metrics do not; this chapter does not attempt to explain why, only to report it.

The defensible reading is narrow and precise:

The correct paired training correspondence is doing substantial work for this supervised ridge bridge. Destroying it collapses counterpart recovery and neighborhood agreement to near-floor levels on this benchmark.

This control must not be read more broadly than that. It is not a demonstration that unpaired translation cannot work, and it is not a lower bound on the performance of a competent unpaired alignment method. Chapter 18 already covered genuine unpaired alignment research — vec2vec and mini-vec2vec — which attempts to infer correspondence structure from the geometry of the two spaces, using adversarial, cycle-consistency, or iterative pseudo-pair-matching machinery specifically built for that purpose. This control does the opposite: it takes the same supervised regression estimator that expects true correspondences and deliberately feeds it a random permutation instead. It measures what happens when supervision is sabotaged, not what a real unpaired method could achieve. Keep those two regimes — sabotaged paired supervision and genuine unpaired correspondence discovery — entirely separate.

Direction changes the profile

Reverse the bridge — Qwen → mxbai, same fixed training anchors, same held-out sentences, same ridge construction:

qwen -> mxbai, paired ridge (n_train = 2,336)

counterpart recovery
  cosine_to_target           0.9246
  recall@1                   0.8862
  recall@10                  1.0000
  mrr                        0.9366

structural fidelity
  agreement@10               0.8154
  local order (Spearman)     0.7232
  triplet agreement          0.9162
  cluster ARI                0.3628

Numerically, in this frozen run: Recall@1 rises from 0.8676 to 0.8862; MRR rises from 0.9221 to 0.9366; agreement@10 rises from 0.7964 to 0.8154; local order rises from 0.6987 to 0.7232; cluster ARI falls, from 0.4750 to 0.3628. No uncertainty analysis exists for either direction — report these as point-estimate differences in this one run, not as statistically reliable or significant differences, since no bootstrap or repeated-split evidence was gathered for either.

The result stands regardless: preservation is directional, and not uniformly so across every metric. A → B and B → A are, per Chapter 20’s discipline, two entirely separate evidence records — each with its own source_space_hash, target_space_hash, direction, and representation path. A→B evidence never authorizes B→A use, and this measured reversal shows precisely why the two need independent evaluation rather than an assumed symmetry: the reverse direction is not simply “the same result, mirrored” — it trades a materially lower cluster-agreement score for a modest gain on every other measured field.

More anchors help — unevenly, and without settling the question of a fixed gap

Wave 6 measures exactly two training-anchor conditions for mxbai → Qwen: n_train = 800 and n_train = 2,336.

                          n=800        n=2,336       change

cosine_to_target          0.8388       0.8837        +0.0449
recall@1                  0.7385       0.8676        +0.1291
recall@10                 0.9983       1.0000        +0.0017
mrr                       0.8454       0.9221        +0.0767
agreement@10               0.7355       0.7964        +0.0609
local order (Spearman)     0.5848       0.6987        +0.1139
triplet agreement           0.9134       0.9197        +0.0063
cluster ARI                 0.4578       0.4750        +0.0172

More paired training anchors improved every measured field in this two-point comparison — but by different, uneven amounts, not one uniform shift. Recall@10 was already nearly saturated at the smaller anchor count (0.9983) and had almost nowhere left to improve; Recall@1 and local order moved by roughly twelve points each; cluster ARI and triplet agreement barely moved at all. Do not compress this into “roughly five points on both contracts” — no such uniform figure describes what actually happened.

Resist a further temptation this comparison invites: computing a single “recovery-minus-structure gap” and asking whether it holds steady across the two anchor counts. No such gap has a single well-defined value, because the two groups contain multiple heterogeneous metrics with no shared scale — recall@1 − agreement@10 moves from 0.0030 at n=800 to 0.0712 at full training size, while recall@10 − agreement@10 moves from 0.2628 down to 0.2036; these two candidate “gaps” tell opposite numerical stories from the same data, which is itself evidence that no single gap quantity is well defined here. What does hold at both measured anchor counts, and is worth stating as the clean, defensible result:

At both n=800 and n=2,336, near-saturated top-10 counterpart recovery (Recall@10 of 0.9983 and 1.0000 respectively) coexists with substantially lower neighborhood agreement (0.7355 and 0.7964). Only two anchor counts were measured — this is not a scaling curve, and no plateau or asymptotic behavior has been established.

Diagnostic regions, not automatic permissions

It is useful to place a bridge’s two profiles on two axes — but only as a way of generating hypotheses about what to test next, never as a substitute for the requirement-plus-evidence authorization Chapter 20 already built:

                              STRUCTURAL FIDELITY
                              lower                    higher

  COUNTERPART      higher     candidate for counterpart   candidate for uses that also
  RECOVERY                    lookup / migration only —   depend on target-space
                               still needs its own          structure, pending its
                               acceptance criterion          own structural evidence

                   lower      likely needs re-evaluation   an unusual combination this
                               before any use                chapter's evidence does not
                                                              illustrate

The upper-left region matters most for this chapter’s headline result. A bridge sitting there is not automatically authorized for anything — it is a candidate whose profile suggests it may be sufficient for a consumer whose stated requirement is counterpart recovery and nothing more, provided that consumer’s own policy, under Chapter 20’s machinery, actually confirms that requirement is satisfied and that no structural property is separately required. Phrase every conclusion this way, not as a direct verdict:

consumer whose only declared requirement is counterpart recovery:
  inspect recall@1 / recall@10 / mrr against that consumer's own stated bar first

consumer relying on neighborhoods, ranking order, or cluster structure:
  must separately specify and check a structural requirement —
  counterpart-recovery evidence alone does not speak to it

Do not write that a bridge is “suitable for lookup” or “unsuitable for clustering” as a direct conclusion of this experiment. Migration itself is not one requirement — locating a paired record, replacing a stored vector, preserving ranking behavior, and preserving a calibrated threshold are four different consumer needs, each requiring its own evidence under Chapter 20’s discipline, and this chapter’s job stops at making that evidence legible, not at issuing the permission.

What this chapter establishes and what it does not

Establishes: counterpart recovery and structural fidelity are distinct properties computed from the same translated vectors; where a structural metric uses a target-neighbor candidate set, Wave 6 excludes the paired counterpart so counterpart recovery cannot mechanically contribute neighborhood-overlap credit. The measured headline profile for mxbai → Qwen paired ridge shows perfect top-10 counterpart recovery (Recall@10 1.0000) and strong but non-perfect top-1 recovery (Recall@1 0.8676, MRR 0.9221), while separate structural measurements report neighborhood agreement 0.7964, local-order Spearman 0.6987, triplet agreement 0.9197, and cluster ARI 0.4750, all on 589 held-out sentences under a deterministic item-ID split. The high-structural-agreement mxbai → bge control shows substantially higher agreement (0.9127) and local order (0.8944), but its cluster ARI (0.4761) is essentially unchanged from the headline pair, so this control supports the discriminative value of some structural measurements here and not others. The shuffled-correspondence control collapses counterpart recovery and neighborhood agreement to near-floor values, showing the true paired correspondence carries substantial signal for this supervised estimator without saying anything about competent unpaired alignment methods. Direction changes the profile non-uniformly across metrics, and more paired training anchors improved every measured field between the two tested anchor counts, unevenly, without establishing a scaling curve or a fixed recovery-structure gap.

Does not establish: that a bridge with strong counterpart recovery is generally unsuitable, or generally suitable, for any named consumer — that decision requires Chapter 20’s requirement-plus-evidence machinery, not this chapter’s measurements alone; that structural fidelity logically follows from, or is logically required by, counterpart recovery in either direction beyond the one measured counterexample; that native structural similarity (CKA) causes or predicts bridge fidelity, beyond two consistent-but-unreplicated observations; that cluster ARI is inherently a noisy or unreliable statistic in general, as opposed to weakly discriminating in this one frozen comparison; that any numerical gap between counterpart-recovery and structural-fidelity metrics is stable, constant, or governed by a scaling law with respect to anchor count; that this benchmark’s designed, templated sentence corpus generalizes to natural production text; or that shuffled-correspondence performance bounds what a genuine unpaired alignment method could achieve.

Lab 22: separate the two contracts, then attack the separation with controls

MEASURED — artifact experiments/embeddings-from-first-principles/wave6/artifacts/cross-space-benchmark.json. REPRODUCIBLE — python embed_corpus.py then python run_wave6.py.

Question. Can a bridge retrieve the right point while reproducing only part of the surrounding structure — and does that gap survive scrutiny from positive and negative controls, direction, and anchor count?

Step 1 — freeze the benchmark provenance. Corpus: n = 2,925, content_sha256 = 0c47690b068f7b07328fabea1fb9aa05af6e067d5b8160792dad588f71b35517, a designed/templated RELATE + RELATE-DOC sentence pool, not a natural corpus. Split: sha1(id) % 5 == 0 → test, giving n_train = 2,336, n_test = 589 — an item-ID split, explicitly distinct from the split_entity entity-family split used in Chapters 18–21.

Step 2 — freeze source and target spaces. Source: mxbai-embed-large-v1, 1024-d. Target: Qwen3-Embedding-8B, 4096-d. All embeddings L2-normalized before fitting and scoring.

Step 3 — fit the exact headline map. Ridge(alpha=1.0), default fitted intercept, trained on L2-normalized paired training rows, predictions L2-normalized.

Step 4 — measure counterpart recovery, defining each field precisely: cosine_to_target = 0.8837 (mean cosine to the paired native target); recall@1 = 0.8676 (paired target is the single top-ranked candidate); recall@10 = 1.0000 (paired target appears anywhere in the top ten); mrr = 0.9221 (mean reciprocal rank of the paired target).

Step 5 — measure structural fidelity on the same translated test vectors, with the paired counterpart explicitly excluded before any neighborhood computation: agreement@10 = 0.7964; local order (Spearman) = 0.6987 — a correlation coefficient, not a percentage; triplet agreement = 0.9197 — a separate, globally sampled pairwise-order statistic, distinct from local order; cluster ARI = 0.4750 — a chance-adjusted clustering-agreement statistic, not a percentage of structure preserved.

Step 6 — state the counterexample precisely. Not “the bridge failed.” Precisely: recall@10 = 1.0000 does not imply agreement@10 = 1.0000 or local order = 1.0000. This is what the experiment establishes.

Step 7 — run the high-structural-agreement positive control. mxbai → bge-large, native CKA 0.9844: recall@1 = 0.9932, recall@10 = 1.0000, mrr = 0.9966, agreement@10 = 0.9127, local order = 0.8944, triplet agreement = 0.9722, cluster ARI = 0.4761. Interpret precisely: agreement and local order are substantially higher than the headline pair; cluster ARI is essentially unchanged (0.4761 vs 0.4750) — so this control validates some structural diagnostics as discriminating and leaves cluster ARI’s discriminating power for this comparison unconfirmed.

Step 8 — run the shuffled-correspondence negative control. Same ridge construction, training target rows randomly permuted: recall@1 = 0.0034, recall@10 = 0.0102, mrr = 0.0111, agreement@10 = 0.0421, local order = -0.1879, cluster ARI = 0.0729, triplet agreement = 0.6485 (notably still well above the near-floor values of the other fields). State: destroying pair correspondence collapses counterpart recovery and neighborhood agreement, showing correct correspondence is doing real work. Do not read this as a bound on competent unpaired alignment methods, which infer correspondence rather than have it deliberately destroyed.

Step 9 — check directionality. Qwen → mxbai: recall@1 = 0.8862, recall@10 = 1.0000, mrr = 0.9366, agreement@10 = 0.8154, local order = 0.7232, cluster ARI = 0.3628. Report only as point-estimate differences from the forward direction in this one run — recovery and local order both rise, cluster ARI falls — never as statistically established differences.

Step 10 — check the anchor-count comparison. n=800 versus n=2,336, as tabulated above. Every measured field improved with more anchors, by different amounts; Recall@10 was nearly saturated already at n=800. Two anchor counts do not establish a scaling curve, a plateau, or a fixed recovery-structure gap.

Step 11 — write the fidelity observation as evidence, with no verdict. Report counterpart_recovery and structural_fidelity as two separate groups of measured fields, each with its exact definition and reference; do not compute or report a PASS/FAIL, and do not collapse either group into one scalar.

Step 12 (PROPOSED — no artifact backs this) — sensitivity work. Repeated train/test partitions; bootstrap resampling of test items; varying the neighborhood k for agreement@k; varying the local-order window beyond the top-20 native window; varying k and the random seed for the KMeans clustering step; evaluating a natural, non-templated corpus. None of these results currently exist.

Try it yourself

Among held-out objects whose paired target remains rank 1, find one with a large neighborhood disagreement and inspect exactly which native neighbors moved out and which translated neighbors moved in. If an object with four or more replacements exists, use it; otherwise use the largest disagreement you actually observe rather than assuming one. Then compare the aggregate forward- and reverse-direction fidelity reports: local order can be inspected per object, but cluster ARI is a whole-test-matrix statistic and should not be attributed to one example. Do not average heterogeneous metrics into one score before you have named which one your downstream use actually depends on.

Companion component: the fidelity observation

This is an evidence artifact, not a policy artifact — it stays entirely on Chapter 21’s side of the evidence/authorization boundary, with nothing resembling a verdict, a consumer requirement, or a usable_for field:

bridge_fidelity_observation:
  bridge_id:
  bridge_version:

  identities:
    source_space_hash:
    target_space_hash:
    derived_output_space_hash:
    direction:                     A_to_B

  transformation:
    family:                        ridge
    alpha:                         1.0
    fit_intercept:                  true
    output_normalization:           l2

  evaluation_contract:
    corpus_hash:                    0c47690b...
    split_rule:                     "sha1(id) % 5 == 0 -> test"
    n_train:                        2336
    n_test:                         589
    corpus_scope:                   designed_templated_book_corpus

  counterpart_recovery:
    cosine_to_target:      { value: 0.8837, definition: "mean cos(T(source_i), native_target_i)" }
    recall_at_1:            { value: 0.8676, definition: "paired target is the top-ranked candidate" }
    recall_at_10:           { value: 1.0000, definition: "paired target within top-10 candidates" }
    mrr:                    { value: 0.9221, definition: "mean reciprocal rank of the paired target" }

  structural_fidelity:
    agreement_at_10:
      value:                 0.7964
      counterpart_excluded:  true
      definition:            "mean |kNN(native_i) ∩ kNN(translated_i)| / 10, native_i counterpart excluded"

        local_order:
      value:                 0.6987
      statistic:              spearman
      native_window:          20
      translated_window:      10
      candidate_set:          "native top-10 UNION (translated top-10 INTERSECT native top-20)"
      definition:             "Spearman of native-target vs translated-point cosine ranks over the selected local candidate set"

    triplet_agreement:
      value:                 0.9197
      definition:             "fraction of sampled (a,b,c) triples whose cosine preference sign matches native vs translated"

    cluster_ari:
      value:                 0.4750
      clustering:
        algorithm:            kmeans
        k:                    25
        n_init:                10
        random_state:          0

  controls:
    high_native_agreement_control_ref:   mxbai->bge-large_ridge   # native CKA 0.9844, not a null map
    shuffled_correspondence_ref:         mxbai->qwen_ridge_shuffled
    reverse_direction_ref:               qwen->mxbai_ridge
    anchor_count_ref:                    mxbai->qwen_ridge_n800

  uncertainty:
    repeated_split:          false
    bootstrap:                false

  task_fidelity_observation_refs:
    # references to OTHER evaluation observations, not new Wave 6 results —
    # e.g. Wave 3's retrieval_ratio, calibration_transfer, relation_profile_corr

  provenance:
    artifact_ref:              wave6/artifacts/cross-space-benchmark.json
    code_hash:

Two design choices are worth naming. First, only the four fields Wave 6 actually measures appear under structural_fidelity; a future agreement_at_50 or a wider local-order window would be a genuine extension of this artifact, not something to list here as though already measured. Second, task_fidelity_observation_refs points outward, to other evaluation observations (Wave 3’s retrieval_ratio, calibration_transfer, relation_profile_corr — Chapter 21’s evidence) rather than duplicating them, keeping each experiment’s provenance in exactly one place.

Embedding Observatory progression

The Observatory should now be able to display, for any bridge, three separated evidence groups rather than one blended profile: counterpart recovery (cosine_to_target, recall@1, recall@10, mrr), structural fidelity (agreement@k, local rank correlation, triplet agreement, cluster comparison), and task fidelity (references out to task-specific observations from earlier chapters). For each field it should show the exact population, candidate-set or whole-matrix construction, and metric definition. Where a target-neighbor candidate set is involved, it should show whether the paired counterpart was excluded; for local order, it should show the native and translated windows and candidate-set construction; for clustering, it should show the KMeans configuration. For every observation it should also show direction, anchor count, which controls exist alongside it (a high-structural-agreement control, a shuffled-correspondence control, a reverse-direction observation), and whether any uncertainty estimate is available — here, none is.

A reader consulting the Observatory for this bridge should be able to see, at a glance and without re-deriving it: this bridge retrieves its paired point extremely well but reproduces the surrounding target neighborhood less faithfully — and should be able to click through to the exact definition behind every number that claim rests on.

Failure modes

  • Calling Recall@10 “geometry preservation.” It measures whether the paired target appears in the top ten under this candidate pool — nothing about the neighborhood around that point.
  • Letting the paired counterpart into a neighborhood-agreement computation. Wave 6 explicitly excludes it; an implementation or explanation that does not will silently inflate structural-fidelity numbers with counterpart-recovery credit.
  • Calling local order 0.6987 “70% of the ordering.” It is a Spearman correlation coefficient, not a percentage of sequence preserved.
  • Calling cluster ARI 0.4750 “47.5% of cluster structure preserved.” ARI is chance-adjusted agreement, not a percentage; it can be near zero or negative.
  • Claiming the high-structural-agreement control validates every structural metric. It strongly separates agreement@10 and local order from the headline pair; it does not separate cluster ARI.
  • Calling the mxbai → bge control a null map or a same-space control. It is a genuinely fitted ridge bridge between two distinct encoder identities, evaluated the same way as the headline bridge.
  • Diagnosing cluster ARI as “noisy” without a sensitivity analysis. No seed sweep, k sweep, or bootstrap was run; the honest statement is that it did not discriminate this one comparison.
  • Building a monotonic staircase from heterogeneous metric magnitudes. Recall, correlation, overlap, and ARI are not positions on one shared fidelity scale.
  • Assuming strong structural fidelity implies strong counterpart recovery, or the reverse in general. Only the measured direction — excellent recovery without excellent fidelity — is established here.
  • Saying more anchors leave “the gap” unchanged. No single well-defined gap exists across these heterogeneous metrics, and only two anchor counts were measured.
  • Treating the shuffled-correspondence control as evidence against unpaired translation. It sabotages a supervised estimator’s training labels; it says nothing about methods built to infer correspondence.
  • Treating two CKA-labeled pairs as a scaling curve or causal law. Two observations are consistent with a pattern; they do not establish one.
  • Calling counterpart recovery “identity recovery.” Chapter 17 already reserved “identity” for space_hash; use “counterpart” or “paired-target” recovery instead.
  • Putting PASS/FAIL, verdict, or consumer_requires inside the fidelity observation. That belongs to Chapter 20’s separate authorization layer, never to this chapter’s measurement artifact.

What this chapter established

  • Chapter 21 showed that a metric’s name can hide what it computes. This chapter shows that the word “retrieval” itself hides at least three contracts — counterpart recovery, structural fidelity, and task retrieval quality — none of which substitutes for another.
  • Counterpart recovery asks whether T(E_A(x)) recovers its own paired E_B(x). Structural fidelity asks whether it reproduces the relationships surrounding that counterpart — with the counterpart explicitly excluded from the computation, so strong recovery cannot inflate it.
  • On Wave 6, the same translated vectors deliver perfect top-10 counterpart recovery (1.0000) alongside agreement@10 0.7964 and local-order Spearman 0.6987. Finding the point and rebuilding its surroundings are different achievements — and those heterogeneous statistics should be read by their own definitions, never collapsed into one structural-fidelity percentage.
  • The controls audit the metrics, not only the system. A high-structural-agreement control lifts agreement and local order substantially while leaving cluster ARI essentially unchanged, which leaves ARI’s discriminating power unconfirmed for this comparison. A shuffled-correspondence control collapses recovery to the floor — showing the paired signal does real work, without bounding what a genuine unpaired method could achieve.
  • Direction changes the profile non-uniformly, and more anchors improved every field unevenly across only two measured counts — so no scaling curve, plateau, or fixed recovery-structure gap was established. The fidelity observation records both contracts with exact definitions and no PASS/FAIL; authorization remains Chapter 20’s responsibility.

Next

Wave 6 has shown that putting a translated point near its target is not enough to reconstruct the target’s relationships. The bridge was trained by regression toward paired target vectors — pointwise reconstruction was the optimization target, not Recall@1, not agreement@10, not cluster ARI, which are evaluation metrics computed afterward, never the loss the map was actually fit against. That gap between what was optimized and what was measured is exactly the opening for the next question. It is not simply “how do we preserve geometry?” It is: whose geometry — the source’s, the target’s, or the downstream consumer’s task — should have authority over the translation, and can a training objective be built to preserve it deliberately rather than as a side effect of chasing one point at a time?