← Embeddings From First Principles

Can One Embedding Space Be Translated Into Another?

Suppose two models represent the same objects. Can we learn a map T such that T(E_A(x)) ≈ E_B(x)? Fit the simplest plausible transformation, test it on entity families it never saw, and discover that ≈ is not one number — it is a preservation contract that different properties satisfy at radically different levels.

Part VI — Crossing Embedding Spaces · Can a map be learned at all?

The question, stated carefully

We have two encoders, E_A and E_B, and a set of objects x. Each object has two representations: E_A(x) in space A, E_B(x) in space B. Chapter 16 established that these are distinct declared spaces with no coordinate correspondence assumed by default. When dimensions permit, raw cross-space arithmetic is numerically possible, but without an alignment contract it has no authorized cross-space semantics. That is not the same claim as “the spaces are unrelated.” Chapter 16 also measured, between several genuinely independent encoders, substantial neighborhood overlap and high linear CKA — real structural agreement, existing side by side with the absence of any authorized coordinate correspondence. Those two facts coexisting is exactly what makes this chapter’s question worth asking rather than settled in advance.

The question:

Does there exist a map T such that T(E_A(x)) ≈ E_B(x) for objects x the map was never trained on — and by what standard do we judge “≈”?

If such a map exists and is properly measured, it opens a door Chapter 17 deliberately left closed: using a vector produced by model A somewhere model B is expected, without re-running model B on the original text. But existence of a fitted map is not the same as authorization to use it. Chapter 17 already established that a bridge’s output is a derived representation — its own declared identity, its own space_hash, traceable back to its parent but never silently equal to it. Even a map explicitly trained to land in B’s coordinate system, producing output of exactly B’s dimensionality, does not thereby produce a native E_B(x):

native target vector:      E_B(x)
bridged / derived vector:  T(E_A(x))

The two may become interoperable for a specific, measured operation. They remain distinct by provenance regardless. This chapter’s job is to fit the simplest such map and build the measurement discipline that has to sit between “a transformation exists” and “the transformation may be used” — not to close that gap by assumption.

Notice the question has two halves. The first asks whether translation generalizes to objects the map never saw. The second asks what counts as close enough. This chapter argues the second half is load-bearing: exact coordinate equality is not the only thing “≈” could mean, and a map can satisfy one downstream standard while failing another entirely.

Why a simple map might work at all

Chapter 16 gives the empirical motivation, and it is worth stating precisely rather than reaching for an assumption about training data this book cannot support. Nothing here establishes that the two models were trained on overlapping corpora, share an objective, or share an architecture — that information was not measured and should not be asserted. What Chapter 16 did measure is narrower and still useful: independently trained encoders can organize the same item population with substantial structural agreement — high neighbor overlap, high linear CKA — despite having no established coordinate correspondence. That gives a concrete, testable hypothesis: perhaps enough of the relationship between these two spaces is well-approximated by something as simple as a linear transformation, for the specific properties an application needs.

That is a hypothesis, not a proof. Two independently trained representations could differ by a rotation, a scale, a shear, a lossy projection, a genuinely nonlinear deformation, or some combination — including differences no single global map of any family could capture, if the two models simply emphasize different task-relevant information. This chapter tests the simplest candidate first, precisely because it is cheap to fit and easy to falsify. Chapter 19 widens the search to other map families once this one’s boundary is on the table.

As a linear-algebra intuition only — not a claim about what real encoders do — imagine one distributed pattern is represented by one mixture of coordinates in space A and by a different mixture of coordinates in space B. A matrix can recombine coordinates to turn one mixture into the other without any single coordinate in either space needing to mean anything on its own. Chapter 5 already established that individual coordinates are not, in general, meaningful semantic axes; this metaphor illustrates linear recombination, not a retraction of that lesson.

There is genuine prior evidence in the neighborhood, worth citing for what it actually shows. Cross-lingual word embeddings can be aligned with a linear map fit on a bilingual dictionary (Mikolov, Le & Sutskever, 2013), and constraining that map to be orthogonal — a pure rotation — generalizes better and has a closed-form solution (Smith et al., 2017). That is precedent for “try a linear map before a nonlinear one,” not a guarantee that it works for any particular pair of sentence encoders. And Chapter 16’s high measured CKA values are not the same evidence as “a linear bridge will succeed”: CKA is invariant to orthogonal transformation and isotropic scaling, not to arbitrary invertible linear maps, so a high CKA score motivates trying alignment without proving that a fitted linear W will reconstruct target coordinates well. The motivation to try is real. The result still has to be measured.

The simplest map: least-squares linear projection

Collect anchor pairs: objects embedded in both spaces. These are your Rosetta Stone — the same texts, paid for twice, once in each encoder.

X_A  =  [ E_A(x₁); E_A(x₂); ... ; E_A(xₙ) ]     (n × d_A)
X_B  =  [ E_B(x₁); E_B(x₂); ... ; E_B(xₙ) ]     (n × d_B)

The centered, no-intercept form of ridge-regularized least squares is a useful way to build intuition for what this map does:

W = (X_Aᵀ X_A + λI)⁻¹ X_Aᵀ X_B          # ridge-regularized least squares, centered form

T(v) = v W handles the dimension mismatch directly — W is d_A × d_B, so a 384-d vector becomes a 768-d vector by linear recombination, with no padding required. The λI term keeps the solution stable when anchors are few or collinear. Nothing special has to be invented to equalize the widths, but the rectangular map still has to be fit and evaluated.

That equation is intuition, not the exact measured implementation. The Wave harness supplies L2-normalized source and target embeddings, and the measured bridge fits scikit-learn’s Ridge(alpha=1.0) directly on those loaded anchor matrices. Scikit-learn’s default fit includes a fitted intercept — an affine map, not a strictly linear one through the origin — and the bridge then L2-normalizes each predicted target vector before it is used downstream:

T_raw(x) = x W + b            # affine ridge fit, alpha = 1.0, intercept fitted
T(x)     = normalize(T_raw(x))

Fitting this map is easy. Knowing whether it means anything is the hard part — which is why the evaluation protocol, not the formula, is the chapter’s real contribution.

The test that matters: held-out entity families

Fitting W on anchors and evaluating on those same anchors measures how well the map fits its training data — nothing more. A constrained linear map, especially a ridge-regularized one, does not necessarily fit arbitrary anchor correspondences exactly; ridge regularization explicitly resists exact interpolation. But training-set fit still does not establish generalization, and a more expressive map family can overfit the anchor correspondences more severely than this one does. Either way, the only honest measurement of translation quality comes from objects the map never saw during fitting.

RELATE’s split_entity partition makes this stronger than an arbitrary held-out sample of individual rows: every item sharing an entity-family assignment goes to the same train/dev/test bucket, so a test family’s items do not appear among the training anchors. That is meaningfully more informative than a random point split, though lexical patterns, templates, or other structure can still be shared across families. It is held-out entity-family generalization within one synthetic, controlled item population, not a claim about broad real-world distribution shift.

  1. Split anchors into train / test by entity family, so no test entity’s material leaks into training.
  2. Fit T on train only and freeze it.
  3. Evaluate on the held-out test entities, along several distinct properties — not one.
  4. And the downstream question: if you retrieve in space B using T(E_A(query)) as the query vector, how much of native space B’s task-level retrieval quality does that recover?
    flowchart TD
    A["anchor pairs: objects embedded in BOTH spaces"] --> SP["split anchors by entity family: train / test"]
    SP --> FIT["fit T on train — affine ridge, alpha=1.0, output L2-normalized"]
    FIT --> EV["evaluate on HELD-OUT test entity families only"]
    EV --> M1["target-vector reconstruction — cos(T(E_A x), E_B x)"]
    EV --> M2["neighbors stay neighbors — 10-NN overlap vs native B, within test"]
    EV --> M3["task retrieval preservation — bridged nDCG@10 / native nDCG@10"]
    EV --> M4["fine distinctions survive? — structured hard-negative margin ratio"]
    M1 --> STD["define '≈' from downstream need BEFORE fitting: neighbors / rankings / clusters / thresholds"]
    M2 --> STD
    M3 --> STD
    M4 --> STD
  

The discipline: choose which property you need before fitting. A pass on one does not imply a pass on the others — that is the entire finding this chapter is about to walk through. (Generic diagnostics such as target-vector top-1 or top-10 identity-matching are reasonable things one could measure; they are not part of the eight properties the artifact below actually persists, and this chapter reports only what was measured.)

Demonstration: linear translation on RELATE

MEASURED on RELATE v0.1, Wave 3 row 3.4 — artifact experiments/embeddings-from-first-principles/wave3/artifacts/ladder-8property-matrix.json. Affine ridge bridge (Ridge(alpha=1.0), fitted intercept, L2-normalized output), all-MiniLM-L6-v2 (384-d) → all-mpnet-base-v2 (768-d), fitted on split_entity:train anchors (at most 712 in this corpus), evaluated on held-out split_entity:test entity families.

This pair is worth describing precisely rather than reaching for a label that oversells the contrast. The Wave 3 harness records minilm-l6 and mpnet-base as the same broad Sentence-Transformers family, with different backbones and different output widths — 384-d against 768-d. They are not two versions of one model, and they are not a cross-family pair; they are two distinct checkpoints from a shared lineage, at mismatched dimensionality, which is exactly why a linear map handling that mismatch is the interesting first test.

Wave 3 measures eight distinct preservation properties for this bridge, each with its own precise definition — not one universal “bridge quality” score:

property (on held-out split_entity:test)                          measured

target-vector reconstruction — mean cos(T(x), native y)              0.5860
10-NN neighborhood overlap vs native B, within held-out test set     0.7128
retrieval nDCG@10 ratio — bridged / native, on qualifying queries    0.7987
rank-triplet agreement — sampled triplet order vs native             0.7529
calibration-transfer score (native EER threshold applied in T-space) 0.7882
relation-profile Pearson correlation (per-relation mean cosine)      0.8673
structured hard-negative margin ratio — bridged / native             0.2342
held-out / training-anchor reconstruction ratio                      0.6480

Each of these needs its exact definition stated once, because several of the field names invite a more generous reading than the implementation supports:

Target-vector reconstruction (0.5860) is the mean cosine similarity between each held-out object’s bridged vector and its own native target-space vector. It is not a percentage of coordinates recovered, not an error rate, and not a matching accuracy — it is a mean cosine, on unit-normalized vectors, averaged over the held-out entity families.

Neighborhood overlap (0.7128) builds a 10-nearest-neighbor graph among the bridged held-out vectors and a separate 10-NN graph among the native target vectors for the same held-out items, then averages |N_T(i) ∩ N_B(i)| / 10 across those items. It says that, on average, about 71% of an item’s ten nearest neighbors are shared between the bridged and native neighbor sets — not that every item recovers exactly seven of ten, and not a retrieval Recall@10 figure.

Retrieval preservation (0.7987) is not a measure of how many identical documents the two systems return. It is the ratio of the bridged system’s nDCG@10 to the native target system’s nDCG@10, computed only over the qualifying held-out queries — those whose designated target item and at least one grade-3 positive both fall in the test entity families. 0.7987 means the bridged retrieval system reaches about 80% of native target task quality under that held-out protocol; it says nothing about list-level overlap between the two systems’ actual results.

Rank-triplet agreement (0.7529) samples triplets of held-out items and asks whether the bridged space’s ordering — is item a closer to b than to c? — agrees with the native target space’s ordering on the same triplet. It is a sampled pairwise-order agreement rate, not Spearman or Kendall rank correlation and not a query-level retrieval statistic.

Calibration transfer (0.7882) has the least intuitive construction of the eight. The native target space’s approximate EER threshold is fit on held-out positive/negative pairs; that same numeric threshold is then applied, unchanged, to the bridged space’s scores on the same pairs, producing a translated FAR/FRR error rate. The reported value is max(0, 1 − (translated_error − native_error)) — an artifact-specific score that equals 1 only when the native threshold transfers with zero degradation. 0.7882 is not “79% of the threshold was preserved,” not a cosine similarity, and not a probability; it is this specific error-transfer construction, and it should be introduced that way whenever it is reported.

Relation-profile correlation (0.8673) computes the mean cosine for each relation type (paraphrase, negation, entailment, and so on) separately in the bridged space and in the native target space, then takes the ordinary Pearson correlation between the two resulting profiles across relation types. This is a correlation between two short vectors of mean cosines — not a rank correlation, not “the ordering of relation types is 87% correlated” in any per-item sense, and not a claim about individual pair classification.

Structured hard-negative margin ratio (0.2342) is scoped narrowly and matters the most. For each qualifying held-out query, take the first grade-3 positive and the query’s hard negatives whose construction method is specifically structured_perturbation; compute positive_score − max(structured_negative_scores) in native target space and, separately, the same margin in bridged space; average each; report the ratio of bridged mean margin to native mean margin. 0.2342 means the bridge retained only about 23% of the native target’s mean structured-perturbation hard-negative margin under this protocol — a measurement about one specific class of hard negatives and one specific decision margin, not a statement about every hard negative, every relation, or every downstream application.

Held-out / training-anchor reconstruction ratio (0.6480) compares this same bridge’s coordinate reconstruction on held-out test entities against its coordinate reconstruction on its own training anchors. A value well below 1 says plainly that the map reconstructs unseen entity families materially worse than the anchors it was fit on — direct, measured evidence for why the held-out protocol exists at all, and a reminder that “OOD” here names the entity-family test split relative to training, not open-world production drift.

The pattern the numbers make

Look at the eight values as one profile rather than eight independent facts, because the pattern is the chapter’s actual finding:

We asked whether a simple affine ridge map could translate MiniLM into mpnet-base. The reconstruction cosine came back at 0.586 — mediocre-sounding, if coordinate identity was the bar. But 71% mean neighborhood overlap and about 80% of native retrieval quality survived, and the relation-level cosine profile correlated at 0.867. So far, “it basically works.” Then the structured hard-negative margin ratio came back at 0.234. The word “works” stops meaning anything as a single verdict the moment that number is on the table.

This is not a case where one weak metric drags down an otherwise unified story — the measurements describe different properties, and those properties survive to different degrees. Neighborhood structure, task-level retrieval quality, and the relation-level mean-cosine profile retain substantial signal under this bridge. The structured hard-negative margin is much more heavily reduced: it compares the first grade-3 reference positive with the highest-scoring available structured_perturbation negative for each qualifying query, then compares the bridged and native mean margins. Neither result cancels the other. They answer different questions, which is exactly why a bridge needs a preservation profile rather than one verdict.

A bridge does not work or fail in the abstract. It preserves some properties, under some workload, to some measured degree.

That is likely the single most important sentence in this chapter. It also explains why “coordinate reconstruction is the wrong target” is too absolute a lesson to take away. For an application that only needs coarse neighbor recall, 0.586 reconstruction cosine may be irrelevant next to 0.713 neighborhood overlap. For an application that consumes bridged vectors as exact numeric input to some further learned component, reconstruction quality could matter enormously. The right move is not to declare one property universally correct and the rest noise — it is to name the property the consuming operation actually depends on, and measure that one directly, before deciding anything is authorized.

One narrow, carefully scoped illustration of the same tension, drawn from the artifact’s per-relation detail rather than the aggregate ratio above: on this held-out split, native target mean cosine is 0.817 for paraphrase pairs and 0.784 for negation pairs, a paraphrase-minus-negation difference of +0.033. Under this ridge bridge, the corresponding means are 0.818 and 0.834, a difference of −0.016. The ordering of these two aggregate relation means therefore reverses under the bridge. That is measured and worth noticing, but it is not a claim that every individual negation pair outranks every paraphrase pair, and not a claim that negation is “always misclassified.” Chapter 21 examines relation-level preservation in far more depth and with a dedicated experiment; this observation is a preview of why that chapter is needed, not a substitute for it.

What “≈” should mean

Exact coordinate reconstruction — T(E_A(x)) = E_B(x) on the nose — is not automatically the right target, but it is not universally the wrong one either; the correct answer depends entirely on what the consuming operation does with the vector. What downstream systems actually need is usually one of a handful of distinct properties:

neighbors stay neighbors        (retrieval, dedup, clustering)
rankings stay rankings          (search result order, nDCG-style task quality)
clusters stay clusters          (topic organization)
thresholds stay meaningful      (calibrated accept/reject decisions)
fine margins stay separated     (hard-negative discrimination, typed-relation distinctions)

If the operation is retrieval, neighborhood and task-level retrieval preservation are directly relevant properties; this bridge measures 0.7128 and 0.7987 on those two tests. If the operation relies on transferring a native target-space threshold, the artifact-specific calibration-transfer score of 0.7882 is relevant evidence, but whether that is acceptable requires a predeclared operating requirement rather than an adjective added after the fact. If the operation depends on preserving separation from RELATE’s structured_perturbation negatives, the measured margin ratio is 0.2342; that is the evidence to evaluate for that requirement. If some further learned component consumes the bridged coordinates directly and numerically, reconstruction quality may become load-bearing after all.

T(E_A(x)) ≈ E_B(x) is incomplete until ≈ names the specific property that has to survive for the operation you intend.

Define that standard before fitting the map, from what the consuming operation actually needs — then report only what was measured against it, in the vocabulary each metric actually earns.

Fitting a map and authorizing its use are different steps

This chapter has now produced a fitted transformation and eight pieces of preservation evidence. It has not produced a decision about what that transformation may be used for. Keep the same measurement-versus-policy separation this book has applied everywhere else, one more time:

fit the map          →  a transformation, T
evaluate the map      →  a preservation profile, measured on held-out entities
authorize an operation → a later, scoped policy decision (Chapter 20)

Existence of T does not make bridged vectors usable against a native B index; only measured evidence that T preserves the specific property that operation depends on can begin to justify that — and even then, whether it is sufficient is a policy question this chapter does not answer. Do not write “the map works.” Write “the map preserved X under Y protocol, and did not preserve Z.” That is the level of precision the rest of this book has earned, and this chapter’s own numbers are the clearest possible argument for holding it here too.

Translating without anchors — and the universal-geometry conjecture

Anchor pairs are the easy case: you paid model B to embed the same texts model A already embedded. What if you cannot? You have a database of vectors from an unknown or retired encoder and no way to re-embed the source text.

It is still sometimes possible — but “unpaired” does not mean “data-free.” The unpaired methods discussed here still use samples of embeddings from both spaces; what they remove is the requirement for a known one-to-one correspondence between individual rows.

paired:    (x_A, x_B) are known to represent the same object
unpaired:  samples from space A and space B are both available,
           but no known correspondence between individual rows exists

The cross-lingual community solved a version of this problem: start from a rough guess of the mapping (adversarial training, or a small automatically-mined seed), take the pairs that are mutual nearest neighbors under the current guess as pseudo-anchors, fit Procrustes to those, and iterate (Conneau et al., 2018; Artetxe et al., 2018). Recent work carries this idea to sentence encoders. vec2vec (Jha, Zhang, Shmatikov & Morris, 2025) learns a translator between two embedding spaces with no paired data at all, using a shared latent backbone trained with adversarial, cycle-consistency, and pairwise-distance-preservation losses, and reports successful unpaired translation on the text-embedding model pairs and metrics it evaluates. mini-vec2vec (Dar, 2025) reports a much cheaper alternative on its own evaluated settings — a linear map built from tentative pseudo-pair matching and iterative refinement — recovering much of vec2vec’s result at a fraction of the compute, with reported robustness on the benchmarks it tests.

These results motivate a striking conjecture, stated explicitly as a hypothesis in the vec2vec paper — the Strong Platonic Representation Hypothesis (extending Huh et al., 2024):

neural networks trained with the same objective and modality, with different data and model architectures, converge to a universal latent space such that a translation between their respective representations can be learned without any pairwise correspondence.

Two things worth holding at once, without letting either one erase the other:

  • The cited results are real, and they broaden this chapter’s world. They show unpaired translation succeeding on the specific model pairs and metrics each study evaluates — evidence that paired anchors are not always a hard requirement for some meaningful notion of translation, at least in the settings tested. Note the hypothesis’s own explicit scope: “the same objective and modality.” It does not claim to cover every pair of encoders, and this book has no basis to extend it further than its authors do.
  • The metrics reported in that literature are largely global ones — reconstruction-style cosine to target, top-1 or top-few matching accuracy, and similar aggregate measures. The cited results do not themselves establish the specific, fine-grained properties this book has been building toward: whether negation stays distinct from paraphrase, whether A acquired B stays distinct from B acquired A, whether entailment direction is preserved, or whether a calibrated operating threshold’s false-accept rate survives translation. That is not a claim that no one anywhere has ever measured such things — only that the results cited here do not answer those specific questions, which is exactly the kind of question this book’s own measured bridge, above, was built to ask.

So the question for the rest of Part VI is not “is embedding geometry universal — yes or no.” It is narrower and more useful: when two spaces are aligned well on coarse, global metrics, which finer properties came along, and which did not? This chapter’s own measured profile — strong neighborhood and retrieval preservation next to a hard-negative margin ratio of 0.234 — is already a concrete instance of exactly that split. Chapters 19 through 21 build the tools to examine it further.

A successful global alignment is evidence about coarse geometry. It is not, by itself, evidence that fine relational structure or calibrated operating points transferred — those are separate measurements, with their own separate evidence.

What this chapter establishes and what it does not

Establishes: the translation question, precisely stated, including that a fitted map’s output is a derived representation with its own provenance, never silently equal to a native target vector; why a linear map is a reasonable first hypothesis, motivated by Chapter 16’s measured structural agreement rather than an assumed shared training corpus; the affine-ridge-plus-normalization construction actually measured, with alpha=1.0 and no λ sweep; the held-out entity-family evaluation protocol; that Wave 3 measured eight distinct preservation properties for the MiniLM-L6 → mpnet-base bridge, each with a precise, non-interchangeable definition — including 0.2342 for the structured hard-negative margin ratio, 0.7128 for neighborhood overlap, 0.7987 for retrieval nDCG@10 ratio, and 0.8673 for relation-profile Pearson correlation; that these heterogeneous measurements produce a sharply uneven preservation profile on the same bridge, which is the chapter’s central empirical result; that unpaired translation is possible in the settings the cited literature actually tests, and motivates — without establishing — a universal-geometry conjecture; and that “≈” must be defined by the consuming operation’s actual need, before fitting, not read off a single convenient number afterward.

Does not establish: that linear or affine maps are sufficient in general (Chapter 19 tries other map families on the same pair); that this bridge’s measured profile authorizes any specific downstream operation (Chapter 20 builds that policy layer); that embedding geometry is universal (the cited hypothesis is explicitly scoped to shared objective and modality, and tested on coarse metrics); that a λ sweep, an anchor-count scaling curve, or any confidence interval exists for this measurement (none were run); or that any two encoders can be bridged (some pairs may not be, and this chapter does not test that boundary). It establishes the method and the evaluation discipline: fit on train, freeze, measure a profile of properties on genuinely unseen entities, and report each property in the vocabulary it actually earns.

Lab 18: reproduce the bridge, and its exact preservation boundary

MEASURED — artifact experiments/embeddings-from-first-principles/wave3/artifacts/ladder-8property-matrix.json (ridge rung, minilm-l6 384-d → mpnet-base 768-d, ≤712 split_entity:train anchors, evaluated on held-out split_entity:test entities). REPRODUCIBLE — python run_wave3.py 3.4.

Question. What does the simplest fitted bridge preserve on unseen entity families, and which properties fail first?

Step 1 — freeze source and target identities. Record source_space_hash and target_space_hash (Chapter 17) for minilm-l6 and mpnet-base, along with each model’s valid query/document representation protocol.

Step 2 — load paired item anchors. For the same RELATE items, embed with both encoders: X_source (384-d, minilm-l6), X_target (768-d, mpnet-base). Both are L2-normalized by the Wave harness at load time.

Step 3 — split by entity family, not by row. Use split_entity:train / split_entity:dev / split_entity:test. The measured ridge row fits the fixed alpha=1.0 model on train and performs no hyperparameter selection. If you extend the experiment with a sweep, fit candidate models on train, select the predeclared preservation target on dev, then evaluate the chosen configuration once on test. The measured row uses at most 712 train anchors.

Step 4 — fit the exact measured bridge. Ridge(alpha=1.0).fit(X_source_train, X_target_train), with scikit-learn’s default fitted intercept, then L2-normalize every prediction: T(x) = normalize(Ridge.predict(x)). Do not sweep alpha in this reproduction — the measured artifact does not.

Step 5 — evaluate target-vector reconstruction. Mean cosine between each held-out test item’s bridged vector and its native target vector. Measured: 0.5860.

Step 6 — evaluate neighborhood preservation. 10-NN graphs built independently among bridged and native held-out test vectors; mean |N_T(i) ∩ N_B(i)| / 10. Measured: 0.7128.

Step 7 — evaluate task retrieval preservation. Over the qualifying held-out query subset (target item and at least one grade-3 positive both in the test entity families), compute bridged nDCG@10 / native target nDCG@10. Measured: retrieval_ratio = 0.7987. State plainly: this is a task-quality ratio, not list-level retrieval agreement.

Step 8 — inspect the remaining measured properties, defining each before interpreting it: rank_triplet_agreement = 0.7529 (sampled pairwise-order agreement, not rank correlation); calibration_transfer = 0.7882 (native-threshold error-transfer score, not “percent of threshold preserved”); relation_profile_corr = 0.8673 (Pearson correlation of per-relation mean cosines, not rank correlation); hard_negative_ratio = 0.2342 (structured-perturbation margin ratio against the first grade-3 positive only); ood_vs_id_reconstruction = 0.6480 (held-out reconstruction relative to this same bridge’s own training-anchor reconstruction).

Step 9 — write the observation as a vector, not a Boolean. Do not write bridge_works = true, and do not invent qualitative bands such as weak, moderate, or substantial unless a preservation policy defined those bands before the run. Record the measured profile first:

bridge preservation observation (held-out split_entity:test):
  neighborhood_at10:          0.7128
  retrieval_ratio:            0.7987
  relation_profile_corr:      0.8673
  coordinate_reconstruction:  0.5860
  rank_triplet_agreement:     0.7529
  calibration_transfer:       0.7882   # artifact-specific score
  hard_negative_ratio:        0.2342   # structured_perturbation only
  ood_vs_id_reconstruction:   0.6480   # test/train reconstruction ratio

A later operation-specific policy can decide which of those values are sufficient. The measurement artifact should not make that policy decision implicitly.

Step 10 (PROPOSED — no artifact backs this) — tune alpha correctly. Fit candidate alpha values on train, select the best by a predeclared preservation target measured on dev, then evaluate the chosen model once on test. No such result currently exists for this pair.

Step 11 (PROPOSED — no artifact backs this) — an anchor-count curve. If a larger anchor pool is available — RELATE v0.1’s split_entity:train supplies at most 712, so a curve at n ∈ {100, 500, 2000} would need a larger corpus, not this one — repeat Steps 4–9 at several anchor counts to see how preservation scales. No such curve was run against RELATE.

Step 12 (PROPOSED — no artifact backs this) — uncertainty. Bootstrap over held-out entities, or repeat the entity-family split under several seeds, to obtain a distribution rather than a single point estimate for each property. No such result exists; treat every number in this chapter as a point estimate on one frozen split.

Try it yourself

Fit the same affine-ridge bridge on your own paired anchors, split by entity or object family rather than by row. State, before you fit anything, which properties your downstream use actually depends on — neighborhood preservation, retrieval ratio, calibration transfer, hard-negative margin, or another explicitly defined requirement — and evaluate those directly. If retrieval-oriented properties look strong while the structured hard-negative ratio does not, record that asymmetry rather than collapsing it into “bridge works.” Carry the resulting preservation evidence into Chapter 20’s scoped usability decision. A calibrated accept/reject use must be justified by calibration-specific evidence; it should not be accepted or rejected solely from the hard-negative ratio. Do not tune on your test split, and do not let a strong reconstruction cosine stand in for a property you have not actually measured.

Companion component: the bridge — first fields

Chapter 17 already established that a bridge’s output is a new, derived space with its own identity. This chapter’s artifact records the transformation and the preservation evidence behind it; Chapter 20 will later add the scoped usability decision built on top of this evidence — deliberately not added here.

bridge_v0:
  source_space_hash:            <Ch17>
  target_space_hash:            <Ch17>
  derived_output_space_hash:    <Ch17 — the bridge's own output identity, not target's>

  direction:                    source_to_target      # NOT assumed symmetric or invertible

  method:
    family:                     ridge_regression
    alpha:                      1.0
    fit_intercept:               true
    output_normalization:        l2
    paired:                      true

  training:
    anchor_set_ref:               <RELATE release + corpus_hash>
    split_policy:                 entity_family
    train_entity_families:        <ref>
    test_entity_families:         <ref>
    n_train:                      <=712 in this corpus
    source_representation_protocol: <Ch13>
    target_representation_protocol: <Ch13>

  heldout_preservation:
    coordinate_reconstruction:            0.5860
    neighborhood_at10:                    0.7128
    retrieval_ndcg10_ratio:                0.7987
    rank_triplet_agreement:                0.7529
    calibration_transfer_score:            0.7882
    relation_profile_corr:                 0.8673
    structured_hard_negative_margin_ratio: 0.2342
    test_vs_train_reconstruction_ratio:    0.6480

  status: evaluated       # fitted | evaluated | rejected — not yet authorized for any operation

A few properties of this record are worth calling out because they carry real engineering weight. Direction is explicit and asymmetric: this bridge maps source to target; nothing about a rectangular 384 × 768 ridge fit defines or implies its inverse, and a target_to_source bridge, if ever needed, is a separate directional bridge requiring its own held-out evaluation. The anchor contract is provenance, not a bare count: recording only n_train tells a future reader little about what the anchors covered — the entity families, corpus hash, and representation protocols on both sides are what make the held-out numbers interpretable. The bridge is tied to both space identities, not just one: if either source_space_hash or target_space_hash changes, this record no longer matches the new source/target pair and must not silently authorize it. The system needs a new bridge record for the new identity pair. An existing checkpoint may be evaluated as a candidate under that new contract, or the bridge may be refit; either way, the old preservation evidence does not transfer by assumption.

Failure modes

  • Treating bridged output as native B. T(E_A(x)) is a derived representation with its own identity (Chapter 17); it is not E_B(x) merely because it has B’s dimensionality and was trained to approximate it.
  • Evaluating on training anchors. Measures fit, not generalization; always report held-out entity-family performance.
  • Treating “map exists” as “map is usable.” Fitting T and measuring its preservation profile are prerequisites; authorizing an operation with it is a separate, later policy decision (Chapter 20).
  • Calling retrieval_ratio retrieval agreement. It is a bridged/native nDCG@10 ratio on qualifying held-out queries, not list-level overlap.
  • Calling relation_profile_corr a rank correlation. It is ordinary Pearson correlation across per-relation mean cosines.
  • Reading calibration_transfer = 0.7882 as “79% calibrated.” It is a specific error-transfer construction relative to the native EER threshold, not a percentage or a probability.
  • Generalizing the 0.2342 hard-negative ratio to every hard negative or every relation. It is scoped to structured-perturbation negatives, under the first-grade-3-positive convention, on this held-out split.
  • Calling MiniLM-L6 → mpnet-base “cross-family.” The harness records both as the same broad Sentence-Transformers family, at mismatched width.
  • Claiming a λ sweep informed this result. The measured rung fixes alpha = 1.0, with no dev-set model selection.
  • Assuming any sufficiently parameterized map trivially memorizes its anchors. A constrained, ridge-regularized fit resists exact interpolation by construction; the real risk is unmeasured generalization, not guaranteed memorization.
  • Anchors that don’t cover the space. A bridge fit on one topical slice of anchors offers no evidence about vectors from an unrelated slice.
  • Certifying a bridge from one property. A strong reconstruction or retrieval result does not establish calibration transfer or hard-negative margin preservation; each required property needs its own evidence.
  • Reusing a bridge across a space-identity change on either side. A new source_space_hash or target_space_hash means the old bridge record no longer matches the declared pair. Register and evaluate a bridge for the new pair; refit if required, but never transfer the old preservation evidence by assumption.

What this chapter established

  • A bridge is an explicit, directional, measured transformation between distinct space identities, and its output is a derived representation with its own identity — never silently a native target vector.
  • The measured bridge is affine ridge regression (Ridge(alpha=1.0), fitted intercept, L2-normalized output), fit on paired anchors with no λ sweep, and held out by whole entity families rather than random rows.
  • Eight preservation properties diverge sharply on the same bridge. Broad neighborhood, retrieval, and relation-profile structure transfer substantially better than the structured hard-negative margin, which this bridge preserves only weakly — and each of the eight carries a distinct, non-interchangeable definition that its friendly name does not convey.
  • Fitting a transformation and authorizing its use are different steps. This chapter produces evidence, not a usability verdict.
  • “≈” is defined by the consuming operation’s actual need — neighbors, rankings, clusters, thresholds, or fine margins — not by coordinate identity as a default, and not by whichever measured number happens to look most favorable. The bridge_v0 record binds the transformation to both space identities, its derived output identity, its anchor contract, its direction, and its full measured profile, with no usability field at all.

Next

The simplest map has just shown a split personality: broad structure travels with it far better than a specific fine decision margin does. Chapter 18 established how to judge a bridge — a preservation profile, measured on held-out entities, named property by property. It did not establish which transformation deserves to be judged that way. Is the gap here because ridge regression is too flexible in the wrong directions, too constrained in the right ones, or simply the wrong family of map for this pair? Chapter 19 widens the toolkit — the null map, orthogonal Procrustes, relative representations, a learned nonlinear map — and runs each one through the exact same protocol this chapter just built, to find out whether a different map family closes the gap this one left open, or whether the gap is a property of the two spaces rather than of ridge regression.