Alignment
Hold the anchors, the split, and the evaluation fixed. Change only the alignment method — the null map, orthogonal Procrustes, affine ridge, a small MLP — and discover that no method wins every preservation property. Alignment is a toolbox of inductive biases, not a ladder from weak to strong.
Part VI — Crossing Embedding Spaces · Which map family survives held-out data?
What Chapter 18 left open
Chapter 18 fit the simplest serious bridge — affine ridge regression — from all-MiniLM-L6-v2 into all-mpnet-base-v2, and evaluated it on entity families the map never saw. The result was not a verdict. It was a profile, uneven enough to make one number impossible to trust on its own:
coordinate reconstruction 0.5860
10-NN neighborhood overlap 0.7128
retrieval nDCG@10 ratio 0.7987
rank-triplet agreement 0.7529
calibration-transfer score 0.7882
relation-profile Pearson correlation 0.8673
structured hard-negative margin ratio 0.2342
held-out/train reconstruction ratio 0.6480
The chapter’s lesson was not “ridge works.” It was that ≈ is a preservation contract — a named property, measured on held-out data, not a single number a map either clears or fails. That leaves an open question this chapter exists to answer:
Was that uneven profile a property of ridge regression specifically, or does any alignment method produce a similarly uneven profile — just with the unevenness distributed differently?
The experimental design is almost surgical. Hold everything fixed except the one thing under test:
same source/target pair
same anchor split (split_entity: train / test)
same base embeddings
same held-out test entities
same preservation metrics
change:
the alignment method
Then see what actually moves.
A toolbox of inductive biases, not a ladder
It is tempting to arrange the candidate methods on one line from “weak” to “powerful” — null, then a rotation, then a general affine map, then something nonlinear — and assume each step is strictly more capable than the last. Part of that picture is useful, but only if we separate a mathematical inclusion from an engineering intuition about flexibility.
For direct source-to-target transformations, the first three classes are genuinely nested:
identity / null T(x) = x
⊂
orthogonal transformation T(x) = x R, RᵀR = I — rotation/reflection only
⊂
general affine map T(x) = x W + b — rotation + scale + shear + projection
A learned nonlinear family loosens the functional form further, but the specific two-hidden-layer GELU network measured in this chapter should not be treated as a literal mathematical superset of every affine map simply because it is nonlinear. It is one more-flexible hypothesis class with its own architecture and optimization. That distinction matters because the experiment compares concrete fitted methods, not abstract function classes. Two other methods worth knowing about do not belong on this line at all, because they answer a different question rather than a more flexible version of the same one.
CCA (Canonical Correlation Analysis) does not fit a map from A’s coordinates to B’s coordinates. It finds paired directions — one subspace in A, one in B — chosen to maximize the correlation between projections onto them, and aligns the two spaces only within that shared subspace, discarding directions on either side that have no counterpart on the other. That is a different objective from “predict B’s coordinates from A’s,” not a more expressive version of it. Maximizing correlation is also not the same thing as preserving retrieval rank: two highly correlated shared directions can still reorder near-ties that ridge or Procrustes would have kept in their original order.
Relative representations do not produce a source-to-target map at all. Instead of asking “what does this A-vector look like translated into B’s coordinates,” they re-express both spaces through their similarities to a shared set of corresponding anchor objects: an item’s relative representation is the vector of its cosine similarities to each anchor, computed independently in its own native space. Both sides land in one new, shared, anchor-indexed coordinate system — there is no learned T to freeze and reuse, and the anchor identities and their order are themselves part of the representation contract, not an implementation detail. The cited method reports invariance to latent isometries and rescalings under the assumptions and setups it studies (Moschella et al., 2023) — a real and useful property in that setting, not a claim that any two embedding spaces become universally compatible this way. And while relative representations need no regression labels, they still need known corresponding anchor objects embedded in both spaces — a different use of paired data, not the absence of it.
So the honest map of the territory has two parts:
direct source→target transformations (nested by permitted distortion):
null ⊂ orthogonal ⊂ affine/linear ⊂ nonlinear
shared/re-expressed representation strategies (a different question):
CCA — a maximally correlated shared subspace
relative representations — re-express both spaces via shared anchors, fit nothing
Treat the whole set as a toolbox of inductive biases, not a strict evolutionary ladder. An alignment method is a set of constraints on what the transformation is permitted to change. Orthogonality forbids rescaling and shearing; affine ridge permits both; a network permits far more than either. Which constraint is useful depends entirely on which property the downstream consumer needs preserved — and that is an empirical question this chapter now measures rather than one a ladder position can answer in advance.
Run the null map first — when it is defined
Before fitting anything, ask whether fitting is even necessary. If a candidate source and target space share the same output width, the cheapest possible “bridge” is to do nothing:
T(x) = normalize(x)
This is not a strawman baseline to beat easily — it asks how much of the measured preservation profile survives direct coordinate reuse before we learn anything. A learned map has to justify itself against that control on the property the consuming operation actually requires; losing one null-map column does not make the fit pointless if it materially improves another predeclared requirement.
MEASURED on RELATE v0.1, Wave 3 row 3.3 — artifact
experiments/embeddings-from-first-principles/wave3/artifacts/null-map-baseline.json.
The implementation requires matching widths, so it runs on exactly one of the book’s three measured model pairs — bge-large (1024-d) and mxbai-large (1024-d) — and skips the other two, which are marked SKIPPED — dims differ in the artifact rather than silently padded into comparability:
bge-large vs mxbai-large, T(x) = x, held-out split_entity:test
coordinate reconstruction 0.9809
10-NN neighborhood overlap 0.8818
rank-triplet agreement 0.8784
retrieval nDCG@10 ratio 0.9952
That is a remarkable null baseline under all four measured properties — reconstruction essentially at ceiling, retrieval quality essentially matching native target performance, with no fitting at all. Read it carefully, because two different overreads are both tempting here.
It is not evidence that these two encoders were trained on the same size, backbone, or objective, or that any of those properties caused the result — the harness records this pair’s metadata (same 1024-d width, both described as retrieval-tuned, different creators) descriptively, and this single control measures an outcome for this specific pair, not a causal mechanism. The honest statement is narrower: this particular pair shows unexpectedly strong direct-coordinate agreement under these four measured properties — nothing in the experiment explains why, and nothing licenses extending the explanation to any other pair.
It is not evidence that bge-large and mxbai-large are “the same space.” Chapter 17 made space identity explicit precisely so that a strong empirical result like this one could never collapse two distinct declared identities into one. bge-large and mxbai-large keep two different space_hash values regardless of how well direct coordinate reuse happens to score between them. What this strong null result licenses is narrower and still valuable: a learned map for this pair would have to improve a property that matters enough to justify its added complexity. Reconstruction (0.9809) and retrieval ratio (0.9952) are near ceiling here; neighborhood overlap (0.8818) and rank-triplet agreement (0.8784) are high but still leave visible room to move.
When an identity map is mechanically defined, measure it before learning anything.
The harder case: mismatched widths, mapping unavoidable
The null control cannot run at all for all-mpnet-base-v2 (768-d) → bge-large-en-v1.5 (1024-d) — there is no way to compare a 768-vector against a 1024-vector coordinate-for-coordinate without first deciding how to handle the width mismatch, and that decision is itself part of whatever method gets used. This is the pair Chapter 19’s main bake-off uses, precisely because mapping cannot be skipped as an option here the way it arguably could be for bge-large/mxbai-large.
Three fitted methods are compared, holding the base embeddings, the split_entity:train anchors, and the held-out split_entity:test evaluation fixed throughout. Each method’s own dimensionality handling and preprocessing is treated as part of that method — not factored out as a shared, invisible preprocessing step, because it is not shared: Procrustes needs a width adapter that ridge does not, and the MLP learns its own nonlinear reprojection end to end.
Orthogonal Procrustes. The classical problem, min_R ‖X R − Y‖² subject to RᵀR = I, solved in closed form by the SVD of XᵀY, assumes the two matrices already share a width in the form used here. The measured harness handles the mismatch explicitly rather than leaving it implicit: because the 768-d source is narrower than the 1024-d target, the source is centered using the training-source mean and then zero-padded up to 1024 dimensions before the orthogonal fit runs in that padded space. (Had the source been wider than the target, the harness would instead PCA-reduce it down to the target width before fitting — a lossy step, unlike zero-padding, which is isometric.) The measured Procrustes pipeline is orthogonal alignment plus an explicit width adapter, and that adapter’s behavior is part of what the method’s results describe — not a detail that can be waved away when reporting what “Procrustes” measured here.
Affine ridge. The same estimator Chapter 18 introduced: the Wave harness supplies L2-normalized source and target embeddings, and Ridge(alpha=1.0) is fit directly on those loaded training-anchor matrices with scikit-learn’s default fitted intercept. Its predictions are L2-normalized before use: T(x) = normalize(xW + b). No alpha sweep was run for this rung.
A two-hidden-layer MLP. One specific, small nonlinear architecture — Linear(d_source, 512) → GELU → Linear(512, 512) → GELU → Linear(512, d_target), output L2-normalized, trained for 300 epochs with a combined MSE-plus-cosine loss. This is a concrete, modest network, not a stand-in for “arbitrary nonlinearity” — its failure to win below is evidence about this recipe, not a verdict on nonlinear alignment as a category.
CCA is not measured here. It remains a real and useful method conceptually, and Lab 19 below sets it up as a Try-it extension using the same anchor split and evaluation contract — but no CCA rung exists in the artifacts this chapter draws on. Do not read any table below as containing a CCA result.
First look: reconstruction and neighborhood, three ways
MEASURED on RELATE v0.1, Wave 3 row 3.5 — artifact
experiments/embeddings-from-first-principles/wave3/artifacts/nonlinear-vs-linear-unpaired.json.
That filename is worth pausing on, because it is a second clean instance of a lesson this book keeps returning to: read the implementation, not the label. Despite its name, row 3.5 fits every method on paired anchors — the same split_entity:train correspondences used everywhere else in this chapter. The artifact’s own internal note says so directly: it exists as a paired proxy for a question the unpaired literature asks about a different setting, not as an unpaired experiment in its own right. Call this what it is: a paired comparison of linear and nonlinear map families, on the same three model pairs used throughout Wave 3, each method fit once (Procrustes and ridge, one deterministic fit apiece) or three times (the MLP, at seeds 0, 1, 2, to see how much its result moves with initialization).
For mpnet-base (768-d) → bge-large (1024-d):
map coordinate reconstruction 10-NN overlap@10 fit time (s)
Procrustes 0.3979 0.7341 0.47
ridge (affine) 0.7893 0.6810 0.69
MLP, mean over 3 seeds 0.7430 0.6621 0.91
MLP coordinate-reconstruction standard deviation across seeds 0/1/2: 0.0067
Read that standard deviation for exactly what it is: the spread of the MLP’s coordinate reconstruction score across three initializations. It is not an uncertainty band on neighborhood overlap, not a confidence interval, and not a run-to-run variance estimate for Procrustes or ridge — both of which are deterministic, single fits with no seed to vary in the first place.
The pattern already refuses to collapse into “one method won.” Ridge reconstructs target coordinates best — its cosine to the native target reaches 0.79, roughly double Procrustes’s 0.40. Procrustes preserves neighborhoods best — 0.73 against ridge’s 0.68. The MLP wins neither. Its reconstruction (0.74) sits below ridge; its neighborhood overlap (0.66) sits slightly below both linear alternatives; and in this measured environment it was also the slowest to fit, at roughly 1.9× Procrustes’s time and 1.3× ridge’s — a fact about this run’s environment, not a portable constant.
This dissociation is not unique to one pair. The same qualitative pattern — the MLP failing to lead either reconstruction or neighborhood overlap — repeats across all three model pairs row 3.5 measures:
Procrustes: recon / nbr@10 ridge: recon / nbr@10 MLP (3-seed mean): recon / nbr@10
bge-large → mxbai-large 0.9577 / 0.8818 0.8814 / 0.7674 0.8757 / 0.8150
minilm-l6 → mpnet-base 0.4658 / 0.7434 0.5860 / 0.7128 0.5048 / 0.6699
mpnet-base → bge-large 0.3979 / 0.7341 0.7893 / 0.6810 0.7430 / 0.6621
Across all three pairs, this two-hidden-layer MLP does not lead either coordinate reconstruction or 10-NN neighborhood overlap over the best simpler alternative. That is real, measured evidence — three model pairs, one corpus, one specific MLP recipe, two quality metrics. It is not evidence that nonlinear alignment is generally unhelpful, that this MLP architecture is the best nonlinear method one could try, or that a different architecture, more anchors, more training, or a genuinely nonlinear source/target relationship would not change the outcome. State the finding at the scope it actually has:
Across these three paired RELATE alignment problems, this MLP recipe did not improve either coordinate reconstruction or 10-NN overlap over the best constrained linear alternative. This is evidence about one MLP recipe, three model pairs, one corpus, and two quality metrics — not a theorem about nonlinear alignment in general.
Second look: the full preservation profile refuses to pick a winner too
Reconstruction and neighborhood overlap are only two of the eight properties Chapter 18 introduced. Widening to the full profile, on the same mpnet-base → bge-large pair, makes the earlier “Procrustes preserves structure, ridge reconstructs points” framing look too tidy.
MEASURED on RELATE v0.1, Wave 3 row 3.4 — artifact
experiments/embeddings-from-first-principles/wave3/artifacts/ladder-8property-matrix.json. The MLP row here uses row 3.4’s own single default-seed fit — a different run from row 3.5’s three-seed mean above, and not directly averaged with it.
property (held-out split_entity:test) Procrustes ridge MLP
coordinate_reconstruction 0.3979 0.7893 0.7524
neighborhood_at10 0.7341 0.6810 0.6717
retrieval_ratio 0.9411 0.7677 0.7603
rank_triplet_agreement 0.7333 0.7882 0.7294
calibration_transfer 0.9829 0.8331 0.8282
relation_profile_corr 0.9098 0.8262 0.8920
hard_negative_ratio 0.6087 0.9767 0.9482
ood_vs_id_reconstruction 0.5741 0.8347 0.7777
Five of these carry precise, non-interchangeable definitions established in Chapter 18, worth restating in one line each rather than re-deriving from scratch: retrieval_ratio is bridged nDCG@10 divided by native target nDCG@10 on qualifying held-out queries — a task-quality ratio, not list-level retrieval agreement. relation_profile_corr is ordinary Pearson correlation between per-relation mean-cosine profiles, computed in bridged space and in native target space — not a rank correlation, and not “relation ordering.” calibration_transfer is the artifact-specific score built by applying the native target’s EER threshold unchanged to bridged scores and comparing the resulting error to native error — not a percentage of a threshold preserved. hard_negative_ratio is the ratio of mean structured-perturbation hard-negative margins, bridged over native, under the same first-grade-3-positive convention as Chapter 18 — not a hard-negative agreement rate. ood_vs_id_reconstruction is the held-out-test reconstruction divided by training reconstruction for that same bridge; here “OOD” means the entity-family test split relative to the entity-family train split, not open-world production distribution shift.
Read the table by which method actually leads each column, rather than reaching for one aggregate ranking:
Procrustes leads neighborhood overlap (0.7341), retrieval ratio (0.9411), calibration-transfer score (0.9829), and relation-profile correlation (0.9098).
Ridge leads coordinate reconstruction (0.7893), rank-triplet agreement (0.7882), structured hard-negative margin ratio (0.9767), and held-out/train reconstruction ratio (0.8347).
The MLP leads none of these eight reported properties for this pair — it sits behind at least one of the other two methods on every column, sometimes close (relation-profile correlation, 0.8920, near Procrustes’s 0.9098), sometimes not (retrieval ratio, below both linear alternatives).
This destroys any tidy story where Procrustes owns “structure” and ridge owns “points.” Ridge does not merely reconstruct coordinates better — it also preserves rank-triplet ordering and, most strikingly, retains far more of the structured hard-negative margin (0.9767 against Procrustes’s 0.6087) than the method that otherwise looks like the stronger overall performer. Whatever intuition suggested that an orthogonal map’s tighter constraint would make it the more trustworthy “structure-preserving” choice across the board does not survive contact with the rank-triplet and hard-negative columns.
It is tempting to explain the reconstruction/neighborhood trade-off causally — to say ridge “chased” individual coordinates “at the expense of” neighborhoods. Resist that framing; nothing in this experiment traces a causal mechanism, only a correlation between two scores on one held-out set. What can be said as mechanism rather than narrative: an affine map has degrees of freedom — scale and shear — that an orthogonal transformation is forbidden to use, and those extra degrees of freedom give ridge more room to move individual points closer to their targets in ways that can also perturb the local relative geometry an orthogonal map is constrained to leave alone. That is a plausible account of why the two methods might diverge this way. It is not a demonstrated causal story, and this chapter does not claim to have measured one.
No method dominates the eight-property preservation profile. Procrustes leads four properties on this pair; ridge leads the other four, including structured hard-negative margin and held-out/train reconstruction ratio. Even “preservation” is not one property, and a single winner cannot be read off the table without first naming what the consuming operation requires.
The rule this experiment actually supports
Put the two looks together and the governing principle sharpens into something more defensible than “pick the simplest map” or “pick the most powerful map”:
An alignment method is an inductive bias, not a rank on a ladder. Procrustes forbids most distortions; affine ridge permits scale and shear; a network permits more still. Which constraint is useful depends entirely on which property the downstream consumer needs preserved.
That leads to a selection procedure grounded in what was actually measured, not in a leaderboard aggregate:
name the required preservation property (or properties) from the consuming operation
↓
evaluate every candidate method on held-out data, against that property
↓
discard methods that miss the requirement
↓
among methods that pass:
prefer lower operational and generalization complexity
unless the extra complexity earns a meaningfully measured advantage
Neither half of the older, cruder rules survives this experiment intact. “Most constrained always wins” fails immediately — Procrustes loses on rank-triplet agreement and loses badly on structured hard-negative margin, both properties a real application might need far more than calibration transfer. “Highest reconstruction wins” fails just as fast — ridge’s reconstruction lead does not carry over to neighborhood overlap, retrieval ratio, or calibration transfer. What is left is conditional, and appropriately so: escalate complexity only when a simpler method fails a predeclared requirement that a more complex method measurably satisfies — never as a default next step, and never certified by a single favorable column.
One more discipline belongs here, because the temptation to invent it is real: no bootstrap interval, repeated split, or significance test exists anywhere in these artifacts, apart from the one MLP reconstruction seed-standard-deviation already scoped above. Do not describe any gap in these tables as “statistically significant,” “within noise,” or “reliably better” — those words claim an analysis this chapter did not run. A gap such as 0.7893 versus 0.3979 is a large numerical separation in this frozen run; it still does not acquire a confidence interval or population-level significance claim merely because it is large. Report the tables as point estimates from one held-out entity-family split, with only the explicitly measured MLP reconstruction seed variation added where available.
What the local paired result does, and does not, say about the unpaired literature
Everything measured in this chapter uses paired anchors — known correspondences between objects embedded in both spaces. The external literature on unpaired alignment — recovering a translation without any known correspondence at all — asks a genuinely different question, and this chapter’s paired result cannot settle it.
vec2vec (Jha, Zhang, Shmatikov & Morris, 2025) learns a nonlinear translator between two embedding spaces with no paired data, using a shared latent backbone trained with adversarial, cycle-consistency, and pairwise-distance-preservation losses, and reports successful translation on the model pairs and metrics it evaluates. mini-vec2vec (Dar, 2025) reports a much cheaper alternative on the benchmarks it tests: pseudo-pair discovery followed by a linear map and iterative refinement, matching or exceeding the original approach’s results at a fraction of the computational cost, with strong reported stability.
The defensible inference from that comparison is narrow: on the specific unpaired benchmarks mini-vec2vec evaluates, a nonlinear translator was not necessary to achieve strong alignment. That is real and useful — it says something about what unpaired alignment requires in at least those tested settings. It is not a general causal law that nonlinearity’s only contribution to embedding alignment is helping an optimizer converge rather than adding useful representational capacity, and this chapter does not draw that stronger conclusion. The local experiment above adds a second, independent data point in the same spirit — nonlinearity did not help in a paired setting either, on these three pairs — but a paired result and an unpaired result are different evidence streams, measuring different problems, and neither should be read as an explanation for the other:
LOCAL, this chapter:
paired anchors, three RELATE model pairs
a two-hidden-layer MLP does not lead reconstruction or neighborhood overlap
EXTERNAL, cited literature:
unpaired correspondence discovery
a linear pipeline matches/exceeds a nonlinear one on tested unpaired benchmarks
Our paired experiment says extra nonlinear capacity did not help this local bake-off. The external unpaired literature asks a different question entirely: how hard is it to discover the correspondence in the first place, when the anchors themselves are unknown? Keep the two apart, and let each stand on its own evidence.
Preprocessing is part of the method, not invisible cleanup
Every method above handles the loaded embeddings differently before or during fitting, and that handling is not incidental — it is part of what defines the method being compared. The Wave harness supplies every embedding L2-normalized at load time; beyond that shared starting point, each method makes its own choices:
method_record:
base_embeddings:
normalization: l2 (applied uniformly by the harness at load time)
source_dimension_adapter: none | zero_pad(target_dim) | pca(k, matrix_hash)
centering: method- and shape-dependent
output_normalization: l2
anchor_coordinate_system: native | relative(anchor_set_ref)
For the chapter’s primary mpnet-base → bge-large comparison, Procrustes subtracts the training-source mean, zero-pads the 768-d source to the 1024-d target width, and then fits the orthogonal transform; had the source instead been wider than the target, the harness would PCA-reduce it. Ridge fits directly on the loaded L2-normalized anchor matrices with no separate preprocessing-centering step, while the MLP learns its nonlinear reprojection end to end. On the equal-width bge-large → mxbai-large Procrustes run, the harness does not invoke that mismatch adapter or its centering step. Two runs are not the same pipeline merely because both are labeled “Procrustes” when their dimensionality handling differs; the adapter is part of what was actually measured and belongs in the record.
This chapter does not claim to have tested centering, whitening, or anchor-selection strategy as independent variables — no such ablation exists in these artifacts. Mean-centering is a genuinely common practice in alignment work and a reasonable thing to try; nothing here measures how much it helps on this corpus, so no claim about its typical benefit is made. If a chosen bridge’s pipeline does include a persisted step like centering or PCA, Chapter 17’s discipline applies without exception: that transformation is part of the bridge’s derivation and belongs in its provenance, not folded silently into “preprocessing” as though it left no trace on the representation’s identity.
Reconstruction versus downstream preservation, restated precisely
Chapter 18 already established that exact coordinate reconstruction is not automatically the wrong target — it depends on what the consuming operation does with the vector. This chapter’s numbers make the same point sharper, because the same map now visibly serves two different masters at once:
target-vector reconstruction: is T(E_A(x)) numerically close to E_B(x)?
(this chapter: coordinate_reconstruction)
downstream / behavioral preservation: does the bridged system support
the same decisions as the native target system?
(this chapter: neighborhood, retrieval ratio,
rank-triplet agreement, calibration transfer,
relation-profile correlation, hard-negative margin)
Neither category is intrinsically the right one to optimize or to report. A consumer that feeds bridged vectors directly into some further learned component, numerically, cares about reconstruction in a way neighborhood overlap cannot substitute for. A consumer doing nearest-neighbor retrieval cares about neighborhood and retrieval-ratio preservation far more than exact coordinate placement — a map with mediocre reconstruction but strong neighborhood overlap, on that pairing alone, is a plausible candidate for a retrieval-scoped consumer, if its predeclared acceptance criterion is actually satisfied by the measured neighborhood number. That last clause matters: this chapter measures preservation; it does not authorize any specific use of it. Whether 0.73 neighborhood overlap or 0.94 retrieval ratio clears whatever bar a real application sets is a policy decision this chapter deliberately leaves open — that is Chapter 20’s job, not this one’s.
“Semantic preservation” remains useful shorthand for readers, but every time it appears in this book it should immediately cash out into one of the named, measured properties above — neighborhood, ranking, relation profile, calibration, hard-negative margin. None of these metrics certifies semantic equivalence in any deeper sense; each certifies exactly the narrow, defined thing it measures.
Lab 19: reproduce the bake-off, and its exact boundary
MEASURED — artifacts
wave3/artifacts/null-map-baseline.json(row 3.3),nonlinear-vs-linear-unpaired.json(row 3.5),ladder-8property-matrix.json(row 3.4). REPRODUCIBLE —python run_wave3.py 3.3 3.4 3.5.
Question. Does changing the alignment method change which properties survive — or just relabel the same trade-off?
Step 1 — choose the primary pair and freeze the contract. source = mpnet-base (768-d), target = bge-large (1024-d). Training anchors: split_entity:train. Evaluation: held-out split_entity:test. Use each model’s own valid query/document representation protocol (Chapter 13) throughout.
Step 2 — confirm the null map is unavailable for this pair, and run it on the pair where it is defined instead. The literal identity map cannot compare 768 coordinates against 1024. Do not pad it and still call the result “null” — that would be a different, undocumented method wearing the null map’s name. Run the null control on bge-large → mxbai-large instead, as its own separate experiment: 0.9809 reconstruction, 0.8818 neighborhood overlap, 0.8784 rank-triplet agreement, 0.9952 retrieval ratio.
Step 3 — fit the Procrustes pipeline, and record its dimensionality adapter as part of the method. For mpnet-base → bge-large: center the 768-d source, zero-pad to 1024, fit the orthogonal transform in that padded space, L2-normalize the output. Evaluate on held-out test entities.
Step 4 — fit affine ridge. Ridge(alpha=1.0).fit(X_source_train, X_target_train), default fitted intercept, L2-normalized predictions. No alpha sweep.
Step 5 — fit the measured MLP, at three seeds. Two hidden layers of width 512, GELU activations, 300 epochs, MSE-plus-cosine loss, seeds 0, 1, 2. Report the mean coordinate reconstruction and neighborhood overlap across seeds, plus seed_std_reconstruction — and only that quantity as an uncertainty measurement; nothing else in this lab has a stored uncertainty estimate.
Step 6 — reproduce row 3.5’s first-look table.
reconstruction nbr@10 fit seconds
Procrustes 0.3979 0.7341 0.47
ridge 0.7893 0.6810 0.69
MLP, 3-seed mean 0.7430 0.6621 0.91
MLP reconstruction seed std: 0.0067
Interpretation: ridge leads reconstruction; Procrustes leads neighborhood overlap; the MLP leads neither. Report no significance claim — none of these gaps has an uncertainty estimate behind it except the MLP figure named above.
Step 7 — reproduce row 3.4’s broader preservation profile, using the single-default-seed MLP run this row actually stores (not row 3.5’s three-seed mean — the two are different runs, and should never be silently merged into one row):
Procrustes ridge MLP (row 3.4, default seed)
retrieval_ratio 0.9411 0.7677 0.7603
rank_triplet_agreement 0.7333 0.7882 0.7294
calibration_transfer 0.9829 0.8331 0.8282
relation_profile_corr 0.9098 0.8262 0.8920
hard_negative_ratio 0.6087 0.9767 0.9482
ood_vs_id_reconstruction 0.5741 0.8347 0.7777
Before interpreting, restate each metric’s exact definition from Chapter 18: retrieval_ratio is bridged/native nDCG@10, not list agreement; relation_profile_corr is Pearson correlation of per-relation mean cosines, not rank correlation; calibration_transfer is the artifact-specific error-transfer score, not a percentage of threshold preserved; hard_negative_ratio is the structured-perturbation mean-margin ratio, not an agreement rate; ood_vs_id_reconstruction is the held-out/train reconstruction ratio under the entity-family split, not a claim about arbitrary production OOD.
Step 8 — state the actual result. There is no universal winner. Procrustes leads neighborhood overlap, retrieval ratio, calibration transfer, and relation-profile correlation. Ridge leads coordinate reconstruction, rank-triplet agreement, structured hard-negative margin ratio, and held-out/train reconstruction ratio. The MLP leads none of the eight reported row 3.4 properties for this pair. The preservation contract — which property the consuming operation actually needs — decides which columns matter, not an aggregate score across all of them.
Step 9 (PROPOSED — no artifact backs this) — a CCA rung. Fit a CCA-based shared subspace using the same split_entity:train anchors and the same held-out evaluation contract as every other rung in this lab. No result currently exists for this or any RELATE pair.
Step 10 (PROPOSED — no artifact backs this) — anchor-selection strategy. Compare an easy/central anchor sample against a relation- or difficulty-balanced sample, holding method and anchor count fixed as far as possible. No such comparison exists in rows 3.3, 3.4, or 3.5 — nothing in this chapter’s evidence base says whether anchor coverage matters more or less than method choice, and no claim about that comparison should be repeated from this lab until it is actually run.
Step 11 (PROPOSED — no artifact backs this) — uncertainty. Bootstrap over held-out entity families, or repeat the entity-family split under several seeds, for every method and every property — not only the MLP’s reconstruction score. No such result currently exists; every number in Steps 6–8 is a single point estimate on one frozen split.
Try it yourself
Run the null map first whenever the source and target widths match — if it already scores near ceiling, a fitted method has to justify its complexity against that control, not against zero. On your own mismatched-width pair, fit Procrustes (recording its exact width adapter), affine ridge, and — only if the linear methods fail a requirement you named in advance — a small nonlinear map. Before comparing any results, write down which one or two properties your downstream consumer actually needs: reconstruction, neighborhood, retrieval ratio, calibration transfer, rank-triplet agreement, or hard-negative margin. Select the least complex method that clears your stated requirement on held-out data — not the method with the single best column, and not the simplest method by default regardless of whether it clears the bar.
Companion component: bridge method comparison
Chapter 18 introduced bridge_v0 — a fitted transformation plus its held-out preservation profile, tied to both space identities. This chapter’s contribution is the evidence needed to choose among candidate methods, recorded honestly enough that the choice is reproducible rather than asserted:
bridge_v1:
...bridge_v0 fields...
candidate_methods:
- method_id: procrustes_v1
family: orthogonal_procrustes
dimensionality_adapter: zero_pad(source_dim=768, target_dim=1024) # or pca(...), or none
centering: applied_to_source_before_adapter
output_normalization: l2
seed: n/a # deterministic closed-form fit
- method_id: ridge_v1
family: ridge_regression
alpha: 1.0
fit_intercept: true
output_normalization: l2
seed: n/a # deterministic fit
- method_id: mlp_v1
family: nonlinear_mlp
architecture: 2x512_gelu
epochs: 300
seeds_evaluated: [0, 1, 2]
output_normalization: l2
controls:
null_map:
applicable_to_this_pair: false # dims differ, 768 vs 1024
applicable_pair_ref: bge-large_vs_mxbai-large
observation_ref: row_3.3
comparison:
evaluation_contract_ref: <split_entity:train/test, corpus_hash>
preservation_observation_refs: [row_3.4_procrustes, row_3.4_ridge, row_3.4_mlp]
fit_cost_observation_ref: row_3.5
selection:
required_properties_ref: <not yet defined — application-specific, Ch20>
observed_leaders:
neighborhood_at10: procrustes_v1
retrieval_ratio: procrustes_v1
calibration_transfer: procrustes_v1
relation_profile_corr: procrustes_v1
coordinate_reconstruction: ridge_v1
rank_triplet_agreement: ridge_v1
hard_negative_ratio: ridge_v1
ood_vs_id_reconstruction: ridge_v1
chosen_method: <none — no requirement has been declared yet>
status: measured_not_authorized
Notice what this record deliberately does not contain: a single chosen_because: "most constrained map within noise of best preservation" field. “Within noise” is not a phrase this chapter’s evidence supports — no uncertainty analysis covers most of these gaps — and “best preservation” is not one scalar this table produces. What the record stores instead is which method led which property, under which exact evaluation contract, leaving the actual selection open until a required-properties list exists to select against. status: measured_not_authorized says precisely that: evidence has been gathered, and a decision has not yet been made — a decision that belongs to Chapter 20, not to this one.
Each candidate method also produces its own derived representation identity, per Chapter 17 — a Procrustes-derived vector and a ridge-derived vector are not the same representation merely because both happen to land in B’s 1024-dimensional coordinate space. method_id, its parameters, and its preprocessing together determine the derived space_hash for that method’s output; “all map into B” is not provenance.
Embedding Observatory progression
By the end of this chapter, the Observatory should be able to display, for a given source/target pair, every candidate method that was fit against it — null (where applicable), Procrustes (with its exact dimensionality adapter), ridge, relative representations, an MLP — and for each: its exact preprocessing and parameters, its anchor contract, its full held-out preservation profile, and its fit/run provenance, including which artifact and which seed count produced each number.
It should make visible which properties each method leads, and which properties were not measured for a given method at all. It should distinguish, explicitly, values drawn from row 3.4’s single-seed preservation matrix from values drawn from row 3.5’s three-seed reconstruction/neighborhood comparison — never silently merging the two into one row, the way an incautious summary table could. It should surface random-seed dependence exactly where it was measured (the MLP’s seed_std_reconstruction) and nowhere else.
What it should not do is declare a winner. best_bridge = procrustes is not a fact this chapter’s evidence supports — on the full eight-property matrix for this pair, Procrustes leads four properties and ridge leads four. The Observatory’s job here is to make that trade-off legible, not to resolve it; resolving it requires a stated operation and a stated requirement, which is exactly what Chapter 20 supplies.
Failure modes
- Skipping the null control when it is mechanically defined. A sophisticated fitted method might be solving a problem direct coordinate reuse had already nearly solved —
bge-large/mxbai-large’s null baseline shows this can happen. - Reading a strong null result as “same space.” Chapter 17’s distinct
space_hashidentities do not collapse because direct reuse happens to score well for one pair. - Treating every alignment method as one nested expressiveness ladder. CCA optimizes correlation in a shared subspace; relative representations re-express rather than map. Neither is “linear regression, but more so.”
- Hiding a dimensionality adapter inside “Procrustes.” Zero-padding versus PCA-reducing the source changes what was actually measured; the adapter is part of the method’s identity, not invisible plumbing.
- Calling the measured ridge fit strictly linear. It is affine —
Ridge(alpha=1.0)with a fitted intercept — and was not swept overalpha. - Treating CCA as a measured rung. It appears nowhere in rows 3.3–3.5; every CCA number in this chapter is a Try-it extension, not evidence.
- Selecting a method by reconstruction alone. Ridge’s reconstruction lead on
mpnet-base→bge-largedoes not carry over to neighborhood, retrieval ratio, or calibration transfer, where Procrustes leads instead. - Declaring one method “the preservation winner.” On the full eight-property matrix for this pair, Procrustes leads four properties and ridge leads four, including ridge’s lead on rank-triplet agreement, structured hard-negative margin ratio, and held-out/train reconstruction ratio. The table does not supply one scalar called “preservation.”
- Calling
relation_profile_corr“relation ordering” or a rank correlation. It is ordinary Pearson correlation across per-relation mean-cosine profiles. - Calling
retrieval_ratioretrieval agreement. It is a bridged/native nDCG@10 ratio, not list-level overlap. - Calling
hard_negative_ratiohard-negative agreement. It is a structured-perturbation mean-margin ratio under the first-grade-3-positive convention. - Treating the MLP’s three-seed reconstruction standard deviation as uncertainty on every metric, or on Procrustes and ridge. It covers exactly one quantity, for one method, in one artifact (row 3.5).
- Calling row 3.5 an unpaired experiment. Its filename notwithstanding, every method in it is fit on paired
split_entity:trainanchors. - Claiming an anchor-coverage effect was measured. No easy-versus-coverage-balanced anchor comparison exists in rows 3.3, 3.4, or 3.5; treat it as an open, PROPOSED question.
- Concluding that nonlinearity “only helps optimization, not capacity.” The local paired evidence shows this MLP recipe did not win on this data; the external unpaired literature shows a linear method matching a nonlinear one on its own tested benchmarks. Neither licenses a general causal law about nonlinear capacity, and the two evidence streams should not be blended into one claim.
- Choosing a bridge method before naming the required preservation property. The same method is not best on every property this chapter measured; a choice made before stating the requirement is not a reproducible choice.
What this chapter established
- Chapter 18 established how a bridge must be judged; this chapter varied the map family under that same discipline, holding the pair, the anchor split, and the evaluation contract fixed. Identity, orthogonal, and affine maps genuinely nest by permitted distortion; the measured GELU MLP is one concrete architecture, not a literal superset claim. CCA and relative representations sit off that line entirely — they optimize a different objective and re-express both spaces rather than fitting a native-target map.
- The null map is a control, not a floor: on
bge-large→mxbai-large, doing nothing scores0.9809reconstruction and0.9952retrieval ratio. Its cause was not measured, and it does not collapse the two spaces’ distinct identities. It is unavailable by construction for the mismatched-width pair, which is why that pair anchors the bake-off. - No method dominates the preservation profile. On
mpnet-base→bge-large, Procrustes leads neighborhood overlap, retrieval ratio, calibration transfer, and relation-profile correlation; ridge leads reconstruction, rank-triplet agreement, hard-negative margin, and held-out/train ratio; the MLP leads none of the eight. “Best alignment” is not a single scalar ranking. - Preprocessing is part of the method being compared, not a shared invisible step — Procrustes’s dimensionality adapter, ridge’s fitted intercept, the MLP’s learned reprojection. Two runs labelled “Procrustes” are not the same method if their adapters differ.
- Method selection follows a predeclared preservation requirement evaluated on held-out data, with complexity escalating only when a simpler method measurably fails it and a more flexible one measurably clears it. The bridge record carries per-property leaders and an explicit
measured_not_authorizedstatus, because no required-properties list yet exists to select against.
Next
We now have several fitted maps, each honestly evaluated against the same held-out contract, and no single one of them crowned. A fitted map is still only a candidate — a function plus a profile of evidence, sitting in the Observatory without a decision attached to it. The remaining step is operational, not experimental: package the chosen method, its exact preprocessing, its provenance, its measured limits, and the precise scopes in which a runtime may actually use it. That packaged object — a map plus the record of what it was shown to preserve, for whom, under what conditions, with an explicit boundary on what it must not be used for — is a bridge. Chapter 20 builds it.