← Embeddings From First Principles

The Embedding Bridge

A map is a function. A bridge is a map plus exact identity, derivation provenance, held-out preservation evidence, and an operation-specific authorization contract — so a fitted transformation never silently becomes a permission. Package the measured MiniLM-to-mpnet Procrustes evidence into an auditable ALLOW / DENY / CONDITIONAL / NOT_EVALUATED decision, and build the registry that fails closed by default.

Part VI — Crossing Embedding Spaces · When may a system actually use one?

A map is not a bridge

Chapter 19 produced several fitted candidates — a Procrustes pipeline with its dimensionality adapter, an affine ridge map, a small MLP — each with its own held-out preservation profile, and none of them a declared winner. Chapter 19 ended deliberately unresolved: status: measured_not_authorized. Deploying any one of those transformations as “space A and space B are now compatible” is the mistake this chapter exists to prevent.

A map is a function. A bridge is a map plus exact identity, plus derivation provenance, plus held-out preservation evidence, plus an explicit, operation-specific record of what a runtime is actually permitted to do with it.

What has to travel with a translation so that a system consuming it knows exactly how far to trust it — and for exactly what?

The answer this chapter builds:

A bridge packages an exact transformation, exact source/target/derived-output identities, derivation provenance, held-out preservation evidence, and operation-specific authorization rules. A metric can support a decision. It cannot silently become the decision.

This is the same move the book has made at every layer beneath it, applied one level higher. Chapter 6 separated a detector from the policy that acts on it. Chapter 14 separated a score from the calibrated operating threshold built on top of it. Chapter 15 separated a signal bundle from the routing policy that consumes it. Chapter 17 separated identity from compatibility from usability. Chapter 20’s job is to make that same separation hold at the boundary between two embedding spaces: preservation evidence is not authorization. It is the thing authorization has to cite.

The bridge output is still a derived representation

Before building the artifact, one fact from Chapter 18 has to survive into this chapter without softening: T(E_A(x)) is not E_B(x). Whatever method Chapter 19 fit — Procrustes, ridge, an MLP — its output is a derived representation, with its own identity, traceable back to the source and target spaces it was built from but never silently equal to either. A bridge artifact has to carry three identities, not two:

source_space_hash        the exact declared identity being translated FROM (Ch17)
target_space_hash        the exact declared identity being translated TO, and evaluated against (Ch17)
derived_output_space_hash  the exact identity of what this bridge actually produces (Ch18)

The target hash names the native coordinate system and behavior the bridge was fit and evaluated against. The derived-output hash names the exact representation this specific method, checkpoint, and preprocessing pipeline actually emits. They are not interchangeable, and a bridge record that collapses them into one field is the first step toward a bridged vector silently masquerading as a native one somewhere downstream.

The bridge artifact

bridge:
  bridge_id:
  bridge_version:

  identities:
    source_space_hash:
    target_space_hash:
    derived_output_space_hash:
    direction:                    A_to_B   # bridges are directional (Ch17)

  transformation:
    method_id:                    <procrustes | ridge | mlp | relative_reps | ...>
    transformation_artifact_hash: <the fitted matrix/checkpoint/parameters, hashed>
    preprocessing:                <dimensionality adapter, centering, normalization — Ch19>
    parameters:                   <alpha, architecture, etc., as applicable>

  training_contract:
    corpus_hash:
    anchor_set_ref:
    split_policy:                 entity_family
    n_train:
    source_representation_protocol:   <Ch13>
    target_representation_protocol:   <Ch13>

  evidence:
    preservation_observation_refs:    [ <row 3.4 observation for this method/pair> ]
    evaluation_contract_ref:          <held-out split_entity:test protocol>
    uncertainty_status:               not_measured   # unless stated otherwise

  authorizations:
    - operation:
      scope:
      direction:
      representation_path:
      requirement_ref:
      evidence_ref:
      decision:            ALLOW | CONDITIONAL | DENY | NOT_EVALUATED
      conditions:
      policy_version:

  lifecycle:
    state:                 fitted | evaluated | active | retired
    created_at:
    retired_at:
    supersedes:

Two structural choices here do real work. First, evidence and authorizations are separate top-level sections, deliberately. The evidence section is append-only, empirical, and never changes retroactively — it records what was measured, once, under a stated contract. The authorizations section is a list of decisions, each one citing a specific evidence reference and a specific policy requirement, and each one scoped to one operation and one direction. Nothing in evidence says ALLOW or DENY; only an entry in authorizations does, and it has to say why.

Second, lifecycle.state and an authorization’s decision are different axes. A bridge can be evaluated — meaning the declared evaluation contract produced one or more preservation observations — while still being DENY for one operation and NOT_EVALUATED for another. evaluated does not mean every property every future consumer might care about has been measured. There is no single global approved flag on the bridge as a whole, because approval is never a property of the bridge; it belongs to a specific bridge, operation, scope, direction, and representation path under a stated policy.

usable_for is the point — but a permission needs a requirement, not just a metric

The temptation this chapter exists to block is short-circuiting straight from a preservation number to a yes/no answer:

retrieval_ratio = 0.9148   →   usable_for: retrieval = YES

That step is invalid on its own terms. 0.9148 is a measurement. It carries no bar inside it. Nothing about the number “0.9148” says whether that is good enough — good enough is a property of the application, stated in advance, not read off the measurement after the fact. A permission entry has to name both halves and show that one satisfies the other:

authorization entry:
  operation:            <precise operation, not a task noun>
  scope:                <consumer / workload>
  direction:            A_to_B
  representation_path:  <exact data flow — see below>

  requirement:
    metric:              retrieval_ndcg10_ratio
    criterion:            >= <application bar, stated by the consumer's policy>

  evidence:
    value:                 0.9148
    observation_ref:       <row 3.4, procrustes, minilm-l6 -> mpnet-base>
    evaluation_contract_ref: <held-out split_entity:test protocol>
    uncertainty:            not_measured

  decision:              ALLOW | CONDITIONAL | DENY | NOT_EVALUATED

A permission has to cite both the requirement and the evidence that satisfied it. Neither half alone is a decision.

“Retrieval” alone is not a precise enough operation to authorize, and this matters more than it might first appear. “Retrieval” could mean a source-space query translated and searched against a native target index; a native target query searched against a bridged, translated source corpus; both sides translated; two native indexes queried separately and their rankings fused; a top-1 lookup; a top-10 context window for a downstream generator. These are different data paths through different representations, and — as the demonstration below shows concretely — evidence gathered on one path does not transfer to another merely because both get called “retrieval.” The scope includes the data path, not just the task noun.

Four decision states, not two

A binary allow/deny list quietly conflates two very different situations: a property that was measured and found wanting, and a property that was never measured at all. Those deserve different states, because they carry different information and, eventually, different remediation paths:

ALLOW          measured evidence satisfies a stated requirement
CONDITIONAL    measured evidence satisfies the requirement only within a stated sub-range or under a stated condition
DENY           measured evidence exists and fails the stated requirement
NOT_EVALUATED  a stated requirement exists, but no qualifying evidence exists for it

These four states assume that an operation-specific requirement has already been declared. If no consumer requirement exists yet, there is no authorization decision to record: the chain stops before this state machine begins, and the registry refuses the operation by default because no matching authorization entry exists.

NOT_EVALUATED is therefore not a synonym for DENY. If clustering has a declared requirement but the bridge was never evaluated on the required clustering metric, saying clustering: NO misreports what happened — it implies qualifying evidence ran and failed, when in fact none exists. NOT_EVALUATED says the honest thing: absence of qualifying evidence is not evidence of measured failure. The runtime consequence is still fail-closed: both DENY and NOT_EVALUATED refuse the operation by default, while a missing requirement produces no authorization entry at all.

Demonstration: is the MiniLM→mpnet Procrustes bridge usable?

MEASURED on RELATE v0.1, Wave 3 row 3.4 — artifact experiments/embeddings-from-first-principles/wave3/artifacts/ladder-8property-matrix.json. Procrustes bridge, all-MiniLM-L6-v2 (384-d) → all-mpnet-base-v2 (768-d), fitted on split_entity:train anchors, evaluated on held-out split_entity:test entity families.

Chapter 19 fit this bridge and reported its profile without resolving it. Here it is again, in full, exactly as measured:

coordinate reconstruction                0.4658
10-NN neighborhood overlap               0.7434
retrieval nDCG@10 ratio                  0.9148
rank-triplet agreement                   0.7098
calibration-transfer score               0.7566
relation-profile Pearson correlation     0.8605
structured hard-negative margin ratio    0.4735
held-out/train reconstruction ratio      0.5562

Ask the chapter’s question directly: is this bridge usable? The honest first answer is that the question, as stated, cannot be answered yet — not because the evidence is bad, but because no operation, scope, or requirement has been named. This is the exact moment where the architecture either holds or collapses. Watch what happens if it collapses: someone reads retrieval_ratio = 0.9148, decides that sounds good, and writes usable_for: retrieval = YES. That single step has smuggled a policy decision into a measurement, and it has done so without ever stating what “good” meant.

retrieval_ratio has exact, narrower semantics than “retrieval works.” The implementation embeds each held-out query in source space, passes it through the bridge, and scores the translated query against the native target test-item pool — then divides that nDCG@10 by the nDCG@10 a native target-space query achieves against the same pool:

retrieval_ratio  =  nDCG@10( T(query_source) vs native target test-item pool )
                     -----------------------------------------------------------
                     nDCG@10( query_target       vs native target test-item pool )

0.9148 means the bridged query retained about 91.5% of native-target nDCG@10 under this specific held-out RELATE protocol. It is one representation path: a source-space query, translated, searched against a native target index. It is not retrieval agreement (list overlap between two systems), not Recall@10, and — critically — it neither measures nor authorizes the reverse path: a native target query searched against bridged source documents. Those are different data flows, and only the first one was measured. Naming this precisely is not pedantry; it is exactly the distinction Chapter 20 exists to enforce, because “retrieval” as an unscoped word would have let evidence from one path silently authorize a completely different one.

This is also not evidence about a legacy-corpus migration, however tempting that framing is given Chapter 17’s opening scenario. Chapter 17’s row 3.2 measured that scenario directly — a mixed index of raw legacy and new vectors, queried by a native target query. Row 3.4 measures something else: a bridged query against a clean, native target-side item pool, on held-out entity families the bridge never saw during fitting. Confusing the two experiments would attribute a migration result to a chapter that never ran a migration probe.

The other seven fields carry their own precise, non-interchangeable meanings, inherited directly from Chapters 18 and 19 and worth restating once more here because an authorization decision is exactly the place where a mislabeled metric does real damage:

  • coordinate_reconstruction = 0.4658 — mean cosine between each held-out item’s bridged vector and its native target vector. Not a percentage of coordinates recovered.
  • neighborhood_at10 = 0.7434 — mean 10-NN overlap fraction between bridged and native neighbor sets, within the held-out test population. Not Jaccard, not Recall@10.
  • rank_triplet_agreement = 0.7098 — sampled pairwise-order agreement on held-out triplets. Not Spearman or Kendall rank correlation.
  • calibration_transfer = 0.7566 — the artifact-specific score built by applying the native target’s EER threshold, unchanged, to bridged scores and comparing the resulting error against native error. Not “76% of a threshold was preserved,” not a probability, not a direct FAR or FRR.
  • relation_profile_corr = 0.8605 — ordinary Pearson correlation between per-relation mean-cosine profiles, bridged versus native. Not rank correlation, not “relation ordering,” not relation-classification accuracy.
  • hard_negative_ratio = 0.4735 — the ratio of mean structured-perturbation hard-negative margins, bridged over native, under the first-grade-3-positive convention. Not hard-negative agreement, not hard-negative accuracy, not a statement about every hard negative.
  • ood_vs_id_reconstruction = 0.5562 — the ratio of held-out test reconstruction to this same bridge’s own training-anchor reconstruction, a ratio, not a reconstruction value in its own right. The artifact does not separately persist a training-set reconstruction number, so nothing here supports reporting “training reconstruction = 1.00” — that figure was never measured, and the historical field name’s “OOD” means held-out entity families relative to training entity families, not open-world production drift.

What 0.5562 does support, stated at the precision the measurement actually earns: reconstruction generalizes substantially less well to held-out entity families than it fits the training anchors, for this bridge. That is real evidence of a train/test gap — worth taking seriously — but it is not, on its own, a causal diagnosis of why the gap exists. “The bridge overfits the anchor entities” names one plausible mechanism; the measurement does not isolate it from other candidate explanations (limited anchor coverage, genuine distributional difference between entity families, or something else). State the gap; do not assign it a cause the experiment did not test.

No clustering evaluation exists in this artifact. No dedup or false-accept-rate evaluation exists in this artifact. Neither of those operations appears anywhere in row 3.4’s measured fields. If a consumer asks whether this bridge is usable for clustering, the correct evidence state is not DENY — nothing was measured and found to fail — it is NOT_EVALUATED. The runtime still refuses the operation by default; the record simply reports honestly why.

One illustrative authorization, built explicitly and labeled as policy

To show the mechanism working end to end, construct one hypothetical consumer and one hypothetical bar — clearly and unmistakably not a Wave 3 result:

ILLUSTRATIVE POLICY — bar invented only to demonstrate the mechanism, not a Wave 3 finding.

consumer:              retrieval-context service
operation:              source-query -> native-target-index retrieval
direction:               A_to_B
representation_path:     source query -> Procrustes bridge -> native target item pool

requirement:
  metric:                 retrieval_ndcg10_ratio
  criterion:               >= 0.90   (chosen for this example only)

evidence:
  value:                   0.9148
  observation_ref:         row_3.4 / procrustes / minilm-l6->mpnet-base
  evaluation_contract_ref: held-out split_entity:test

decision:                ALLOW, under this hypothetical policy only

Contrast it immediately with a second, equally illustrative consumer whose requirement this bridge cannot even be checked against:

consumer:              clustering service
operation:              coarse topic clustering over bridged vectors

requirement:
  metric:                 cluster_agreement
  criterion:               (undefined — no such metric exists in this artifact)

evidence:
  observation_ref:         NONE

decision:                NOT_EVALUATED
runtime behavior:        operation denied by default

Nothing in the second block claims clustering failed. It claims, accurately, that clustering was never asked. That is a stronger and more honest teaching example than inventing a cluster_agreement = 0.82 result would have been — it demonstrates exactly the failure mode this chapter is built to prevent, using the actual absence of evidence rather than a fabricated presence of it.

Calibration transfer needs its own requirement, not an inherited one

calibration_transfer = 0.7566 deserves one more pass, because it is the field most likely to be misread as a ready-made policy verdict. It is not. The artifact first derives an approximate EER threshold in native target space, applies that same numeric threshold to bridged scores, computes the bridged FAR/FRR average, and then records max(0, 1 - (bridged_error - native_error)). The resulting 0.7566 is therefore an artifact-specific transfer score — not a threshold value, not a percentage of calibration retained, and not a direct FAR or FRR. It does not, by itself, authorize or forbid reusing a native threshold on bridged output — that requires a stated transfer requirement from whatever consumer wants to make thresholded decisions on bridged vectors, exactly the way Chapter 14 required a calibration contract before any threshold could gate anything, and exactly the way Chapter 17 required a derived representation to earn its own calibration evidence rather than inherit its parent’s.

threshold_transfer:
  evidence:        calibration_transfer = 0.7566   (row 3.4, this bridge)
  requirement:     not yet declared by any consumer
  authorization:   not derivable from evidence alone

Absent a declared requirement, the safe runtime default is explicit and fail-closed:

Deny native-threshold reuse on bridged output until a calibration policy explicitly validates it under its own contract.

That is a policy rule, chosen because it is the conservative default — it is not a claim that 0.7566 failed some universal calibration bar, because no such universal bar exists anywhere in this book’s evidence.

Directionality: A→B does not authorize B→A

Chapter 17 already established that a bridge has a direction. This chapter’s authorization layer makes the consequence concrete: an entry that reads ALLOW: source_query_A -> native_index_B retrieval says nothing whatsoever about the reverse path, source_query_B -> native_index_A. The reverse direction is a separate bridge contract and requires its own evaluation. For a purely orthogonal Procrustes map, a mathematical inverse transformation exists in principle and may be evaluated as the reverse candidate; even then, having an invertible matrix is not the same thing as having operational authorization for the reverse path, because the reverse direction was never evaluated against its own held-out queries and its own target pool. For affine ridge or an MLP, an inverse may not exist in any clean form at all. Either way: reverse use is a separately recorded, separately evaluated bridge direction, not a free consequence of the forward one.

Composition creates a new path, not a multiplied score

It is tempting to reach for a clean rule: chain an A→B bridge to a B→C bridge, and the composite’s preservation is the product of the two legs’ losses. That rule does not hold. Preservation metrics in this book are not built to compose multiplicatively — a cosine reconstruction score, a neighborhood-overlap fraction, a retrieval nDCG ratio, and a calibration-transfer score are each their own construction, computed against their own reference, and nothing about chaining two transformations guarantees their measured values combine arithmetically into a third.

There is a more basic problem underneath the arithmetic one, and it is the more important lesson: the output of an A→B bridge is a derived representation, not automatically native B identity (this chapter’s own opening point, and Chapter 18’s before it). Feeding that derived output into a B→C bridge fit and evaluated on native B vectors is not guaranteed to behave the way native B input would — the interface between the two legs has never been checked. So the correct first-principles treatment of composition is:

bridge_AB  +  bridge_BC
        ↓
candidate composite transformation, A -> C
        ↓
is bridge_BC's expected input actually what bridge_AB's derived output looks like?
        ↓
if so: evaluate the FULL composite path end-to-end, on its own held-out contract
        ↓
produces its own new preservation observation
        ↓
produces its own new authorization decision — inherits nothing from either leg

Composition does not multiply preservation scores. It creates a new derived transformation whose end-to-end behavior must be measured directly, on its own held-out evidence, before any operation through it is authorized.

A separate diagnostic: the ridge round-trip

MEASURED SEPARATE RIDGE ROUND-TRIP DIAGNOSTIC — Wave 3 row 3.9, artifact experiments/embeddings-from-first-principles/wave3/artifacts/roundtrip.json. Independently fitted forward (A→B) and reverse (B→A) affine ridge maps, evaluated round trip on RELATE v0.1.

This is a different experiment from the Procrustes bridge above, on a different set of fitted transformations, and it should not be silently folded into the demonstration’s evidence. It fits a ridge map in each direction independently, then asks: starting from a native A vector, translate to B, translate back to A — how close is the result to where it started?

pair                          single-hop forward cosine   round-trip return-to-source cosine   round-trip 10-NN overlap

bge-large <-> mxbai-large              0.8814                        0.8582                          0.6981
minilm-l6 <-> mpnet-base                0.5860                        0.7940                          0.7678
mpnet-base <-> bge-large                0.7893                        0.7449                          0.7314

Read this table carefully, because its two cosine columns measure against different references and must never be subtracted from each other as though they shared a denominator. single_hop_forward_cosine compares the forward bridge’s output against the native target vector — the same reconstruction quantity used throughout Chapters 18–20. roundtrip_cosine compares the round-tripped A→B→A result against the native source vector it started from — a completely different comparison, run through two independently fitted maps. For minilm-l6 ↔ mpnet-base, the round-trip figure (0.7940) is actually higher than the single-hop figure (0.5860), precisely because they are answering different questions against different targets — a fact that would look like a paradox only if the two numbers were mistakenly treated as sharing a scale.

The safe, narrow lesson this diagnostic supports:

A separately fitted A→B→A ridge round trip does not perfectly reconstruct the source representation or its neighborhoods on any of these three pairs. A round trip is another transformation path, with its own reference and its own reference frame, and it needs its own end-to-end evaluation — never inferred from the forward leg’s score, and never subtracted from it.

This diagnostic does not describe the chapter’s Procrustes bridge — it uses ridge in both directions, fit independently. It is not a reverse-direction evaluation of the MiniLM→mpnet Procrustes bridge above; that reverse evaluation was never run. And it does not establish a universal composition-loss figure of any size. Chapter 21 returns to round trips and composition with the deeper tools needed to interpret them properly; this chapter’s job was only to show that the question exists and cannot be answered by arithmetic on the forward score alone.

A bridge is pinned to exact identities — a new model does not rewrite history

Chapter 17 established that a new space_hash on either side of a bridge means the bridge no longer matches its declared pair. This chapter sharpens what that actually means operationally, because the tempting shorthand — “a new model version invalidates the bridge” — gets the direction of the claim backwards.

Suppose a bridge was correctly fit and evaluated for source_space_hash = hash_A, target_space_hash = hash_B. Later, a new target model is deployed, with a new declared identity hash_B2. The original bridge’s evidence does not become false. It remains exactly what it always was: a correctly measured transformation and preservation profile for the pair (hash_A, hash_B). What has changed is that this evidence simply does not apply to the new pair (hash_A, hash_B2). That pair requires a new bridge record and new qualifying evidence. An existing transformation artifact may be evaluated as a candidate under the new identity contract, or the bridge may need to be refit; the old preservation evidence does not transfer by assumption.

A bridge is pinned to exact identities. A new identity requires a new bridge record or an explicit re-evaluation; the old bridge remains valid evidence only for its original declared pair and scope.

The deployment system may reasonably choose to mark the old bridge retired in its lifecycle state once hash_B leaves active production use — that is a sensible operational policy, worth keeping. But retiring a bridge from active deployment and falsifying its historical evidence are different acts. The evidence record should stay immutable and queryable for provenance and audit long after the bridge itself stops being used, exactly the way an old calibration record (Chapter 14) or an old evaluation observation (Chapter 13) remains a true historical fact about the space it was measured against, even after that space is no longer in production.

It is tempting to treat “these are vectors from a private or unknown encoder, and nobody has the matching model” as a security property. The cited evidence argues against relying on that.

Embeddings can be inverted. In its evaluated setting, an iterative correct-and-re-embed attack recovers 92% of 32-token inputs exactly, given only embedding access and query access to the encoder, and recovers personal information — including full names — from clinical notes (Morris et al., 2023). That figure describes the method’s evaluated models and data, not a universal property of every embedding model.

Translation removes the “I don’t have the matching model” defense. vec2vec (Chapter 18) demonstrates translation between embedding spaces without paired data, without access to the source encoder, and without predefined correspondences, using shared latent-geometry structure alone. Its authors report that an adversary holding nothing but embedding vectors can recover information sufficient for classification and attribute inference on the translated representations.

The defensible practical principle, scoped to what these two cited results actually support:

Treat embedding vectors as sensitive data. Do not rely on encoder obscurity or coordinate-space incompatibility as an access-control boundary.

This is a corollary of the bridge material, not a new subject, and the book does not pursue embedding security further than this sidebar. It belongs to security classification and access-control policy, not to usable_for — a bridge artifact records what operations are authorized on a representation, not who is permitted to hold that representation in the first place.

What this chapter establishes and what it does not

Establishes: the distinction between a fitted map and a bridge; the bridge artifact, carrying exact source, target, and derived-output identities, transformation and preprocessing provenance, training-anchor contract, evidence separated cleanly from authorization, and a lifecycle state independent of any operation’s decision; that usable_for must name a precise operation, direction, and representation path — not a task noun; that a preservation metric carries no built-in pass/fail bar, and any bar used for illustration must be labeled as policy, never as an experimental result; ALLOW, CONDITIONAL, DENY, and NOT_EVALUATED as four genuinely different states, with the last distinguishing “measured and failed” from “never measured” while still denying the operation by default in both cases; the exact MiniLM→mpnet Procrustes preservation profile, correctly named field by field; that retrieval_ratio = 0.9148 describes one specific representation path — a bridged source query against a native target pool — and neither a migration scenario nor generic retrieval agreement; that bridges are directional and do not self-authorize their reverse; that composition produces a new transformation requiring its own end-to-end evaluation rather than a multiplied score; that the ridge round-trip diagnostic (row 3.9) is a separate experiment whose two cosine columns are not directly comparable; and that a bridge’s historical evidence remains valid for its original declared pair even after a newer space is deployed elsewhere.

Does not establish: that this or any bridge is authorized for any specific consumer’s operation (no real application bar exists in these artifacts to check against); that clustering, deduplication, or any operation outside the eight measured row 3.4 properties was evaluated for this bridge; that the reverse MiniLM←mpnet direction was evaluated; that composed bridges perform predictably from their parts; or that any bridge is a permanent substitute for re-embedding. It establishes the architecture: evidence is recorded as immutable observations under stated evaluation contracts; authorization is decided per operation, per scope, per direction, against a stated requirement; and the runtime denies by default wherever that chain is incomplete.

Lab 20: build one auditable authorization, and one honest refusal

MEASURED — artifact experiments/embeddings-from-first-principles/wave3/artifacts/ladder-8property-matrix.json (Procrustes rung, minilm-l6 → mpnet-base); roundtrip.json referenced only as a separate diagnostic. REPRODUCIBLE — python run_wave3.py 3.4 3.9.

Question. Given a fully measured preservation profile and no application requirement yet, what does it actually take to reach an authorization — and what happens when the requirement can’t be checked at all?

Step 1 — identify the exact bridge. Record source_space_hash (MiniLM-L6), target_space_hash (mpnet-base), derived_output_space_hash (this Procrustes method’s own output identity), the transformation artifact (the fitted orthogonal map plus its training-source mean centering and zero-pad dimensionality adapter, per Chapter 19), and the direction (A_to_B).

Step 2 — attach the actual held-out evidence, and nothing else. All eight measured row 3.4 fields, verbatim, as listed in the demonstration above. Do not add a clustering, dedup, Recall@k, or MRR result — none exists for this bridge.

Step 3 — define one consumer and one precise representation path. For example: consumer = retrieval-context service; operation = source-query -> native-target-index retrieval; representation_path = source query -> bridge -> native target item pool. Write the path out explicitly enough that it could not be confused with the reverse direction or with a mixed-index migration scenario.

Step 4 — define the requirement before looking at whether it passes. Either use a symbolic placeholder (retrieval_ndcg10_ratio >= application_bar) and leave the decision formally unresolved, or construct one fully explicit illustrative bar, clearly labeled ILLUSTRATIVE POLICY, as in the demonstration above. Never let the observed value retroactively suggest what the bar should have been.

Step 5 — produce the authorization trace.

authorization_trace:
  request:
    operation:              source-query -> native-target-index retrieval
    source_space_hash:      <minilm-l6>
    target_space_hash:      <mpnet-base>
    consumer_scope:         retrieval-context service

  bridge:
    bridge_id:              <this Procrustes bridge>
    direction:               A_to_B
    derived_output_space_hash: <bridge's own output identity>

  requirement:
    policy_version:          <illustrative, v0>
    metric:                  retrieval_ndcg10_ratio
    criterion:                >= 0.90   (illustrative only)

  evidence:
    observation_ref:          row_3.4 / procrustes / minilm-l6->mpnet-base
    measured_value:            0.9148
    evaluation_contract_ref:   held-out split_entity:test
    uncertainty_status:        not_measured

  decision:                  ALLOW   (under the illustrative policy only)
  reason:                    "0.9148 >= 0.90 under policy v0 (illustrative)"

Step 6 — demonstrate fail-closed behavior on missing evidence. Ask whether the same bridge is usable for clustering. No cluster-agreement metric exists anywhere in row 3.4. Produce:

decision: NOT_EVALUATED
reason:   no qualifying observation exists for this operation
runtime:  operation denied by default

This is the stronger teaching example, because it demonstrates the actual evidentiary gap rather than a fabricated one.

Step 7 — demonstrate scoped calibration. calibration_transfer = 0.7566 is evidence; it is not, by itself, a decision about threshold reuse. Here the evidence exists but no consumer requirement has been declared, so do not manufacture an authorization state at all: there is no matching authorization entry yet, and the registry denies native-threshold reuse by default. NOT_EVALUATED would mean something different — that a requirement exists but no qualifying measurement exists for it. A real deployment still needs to declare a calibration contract for bridged-output scores specifically (Chapter 14) and test it directly, rather than inherit target B’s native threshold behavior.

Step 8 — show directionality holding. The A_to_B ALLOW entry from Step 5 says nothing about B_to_A. Do not construct a reverse authorization from it, and do not treat row 3.9’s independently fitted reverse ridge map as if it evaluated this bridge’s reverse direction — it did not; it is a different transformation entirely.

Step 9 — show lifecycle behavior under a model change. Suppose target_space_hash (mpnet-base) is replaced in production by a new identity, hash_B2. The bridge record built in Steps 1–7 does not change, and its evidence does not become false — it remains correct evidence for (hash_A, hash_B). Mark it retired in lifecycle.state if it is removed from active deployment, and register a new bridge record for (hash_A, hash_B2). Its evidence must be measured under that new identity pair. You may evaluate an existing transformation artifact as a candidate or refit a new one; what is forbidden is carrying the old pair’s preservation evidence forward by assumption.

Step 10 (optional) — inspect the ridge round-trip diagnostic separately. Reproduce row 3.9’s table for all three pairs, and confirm that roundtrip_cosine and single_hop_forward_cosine are not comparable quantities — they reference different endpoints. Keep this diagnostic’s conclusions entirely out of the Procrustes bridge’s authorizations list; it evaluates a different transformation.

Try it yourself

Take one bridge from your own system. Write out its full preservation profile, field by field, using each metric’s exact definition rather than its friendly name. Then, before checking a single number against anything, write down the precise operation, direction, and representation path one real consumer needs, and the bar that consumer’s own policy requires. Only then compute the authorization. If any property your consumer needs was never measured, record NOT_EVALUATED rather than inventing a DENY — and confirm your runtime still refuses the operation by default in both cases.

Companion component: the bridge registry

Chapter 19 fit multiple candidate transformations for the same source/target pair. A registry modeled as a single bridge value per (source_hash, target_hash) pair creates exactly the cardinality mistake this chapter has to avoid, because it would force a single Procrustes-versus-ridge-versus-MLP winner where Chapter 19 deliberately declined to name one. A pair may be a useful lookup key, but it must resolve to zero, one, or several bridge artifacts rather than one privileged bridge.

bridge_registry:
  bridges:
    <bridge_id>:
      source_space_hash:
      target_space_hash:
      derived_output_space_hash:
      direction:
      lifecycle_state:
      authorizations:              [ ... ]

  find_candidates(source_hash, target_hash):
      → zero, one, or several registered bridge artifacts for this pair

  authorize(source_hash, target_hash, operation, scope, direction):
      → among candidates: select a bridge with a matching, active authorization entry
      → return ALLOW | CONDITIONAL | DENY | NOT_EVALUATED, with the evidence trace

  default:
      no candidate bridge, or no matching authorization entry
      → DENY, operation refused

  lifecycle:
      historical evidence is immutable and remains queryable after retirement
      active-deployment eligibility is governed separately by lifecycle_state

find_candidates can legitimately return several bridges for one pair — different methods, checkpoints, anchor populations, or bridge versions — while each bridge can also accumulate versioned authorization decisions without pretending those policies are new transformations. authorize selects among eligible candidates by matching operation, scope, direction, lifecycle state, and policy, not by picking “the best” on some aggregate score no evidence supports computing. The default, with no exception, is refusal: an unregistered identity pair, a bridge pinned to different hashes, or the absence of a matching authorization entry all prevent the raw cross-space operation from proceeding. That fail-closed outcome should not be mislabeled as a measured DENY when the actual reason is missing applicability, evidence, or policy.

Embedding Observatory progression

By the end of this chapter, the Observatory should be able to answer, for any candidate cross-space operation: which exact bridge would produce the derived vector in question, and what are its source, target, and derived-output identities? Which transformation artifact, checkpoint, and preprocessing pipeline does it use? Which anchor and evaluation contract produced its evidence? Which preservation properties were actually measured, and which are absent? What operation, in what direction, through what representation path, is the current consumer requesting? Which policy requirement applies to that exact operation and scope? Which measured observation is being cited to support the decision? Is the result ALLOW, CONDITIONAL, DENY, or NOT_EVALUATED, and why? Is this bridge still active for these exact identities, or has one side drifted? Are there several candidate bridges registered for this pair, and why was this particular one selected over the others?

This closes a cumulative arc the last several chapters have been building toward:

space identity (Ch17)
  → candidate bridge transformations (Ch18–19)
    → preservation observations (Ch18–19, deepened in Ch21)
      → operation-specific policy requirements (Ch20)
        → auditable runtime authorization, evidence attached (Ch20)

Failure modes

  • Shipping a map as “compatibility.” A fitted transformation with no identity record, no evidence, and no scope is not a bridge — it is an unaudited function someone will eventually misuse.
  • Turning a metric into a permission without a requirement. 0.9148 does not intrinsically mean YES; a bar has to exist before the comparison does.
  • Backfitting a bar after seeing the result. A policy’s acceptance criterion has to be declared by the consumer’s own requirements, not reverse-engineered from whichever number the bridge happened to produce.
  • Calling retrieval_ratio retrieval agreement. It is a bridged-query/native-query nDCG@10 ratio against a native target pool — one specific representation path, not list-level overlap.
  • Reading row 3.4 as a corpus-migration experiment. It evaluates a bridged query against a clean native target pool on held-out entities; Chapter 17’s row 3.2 is the actual mixed-index migration probe.
  • Treating 0.5562 as a test reconstruction value. It is the ratio of test to training reconstruction; the artifact never persists a separate training-reconstruction figure, and none should be invented or assumed equal to 1.00.
  • Calling calibration_transfer a percentage of threshold preserved. It is the artifact-specific error-transfer score defined in Chapter 18, not a direct calibration percentage.
  • Calling hard_negative_ratio hard-negative agreement. It is a structured-perturbation mean-margin ratio, scoped to the first-grade-3-positive convention.
  • Inventing clustering or deduplication evidence that was never measured. The correct state is NOT_EVALUATED, not a fabricated pass or fail.
  • Authorizing one representation path with evidence gathered on a different one. Source-query-to-target-index evidence does not authorize target-query-to-bridged-source-index, even though both are called “retrieval.”
  • Inverting a bridge without re-evaluating it. Direction is part of the contract; a mathematical inverse, where one exists, is not the same as operational authorization for the reverse path.
  • Assuming bridge composition multiplies preservation scores. No such law holds; a composed path needs its own end-to-end evaluation, and needs its interface checked first.
  • Treating a new target-model deployment as falsifying the old bridge’s evidence. The old bridge remains correct evidence for its original declared pair. A new identity pair needs a new bridge record and qualifying evidence; an existing transformation may be evaluated as a candidate, but the old pair’s evidence does not transfer by assumption.
  • Assuming only one bridge can exist per source/target pair. Chapter 19 produced several candidates; the registry has to hold, and choose among, more than one.
  • Treating a bridged vector as a native target vector by identity. It keeps its own derived provenance regardless of how well it performs.
  • Treating encoder obscurity or space incompatibility as access control. Embeddings remain sensitive data whether or not the encoder that produced them is known.

What this chapter established

  • A map is a function. A bridge packages exact source, target, and derived-output identities, transformation and preprocessing provenance, a training-anchor contract, evidence kept structurally separate from authorization, and a lifecycle state independent of any single operation’s decision.
  • usable_for requires a precise operation, direction, and representation path — not a task noun — and a permission entry must cite both the requirement it checks and the evidence that satisfies it. Neither half is a decision alone. Four states — ALLOW, CONDITIONAL, DENY, NOT_EVALUATED — separate measured success, condition-scoped permission, measured failure, and missing evidence; the last two both refuse by default, and where no requirement exists yet, no entry is emitted and the registry still fails closed.
  • Every field in the measured profile names one specific thing. retrieval_ratio = 0.9148 describes a bridged source query against a native target pool on held-out entity families — not generic retrieval agreement, and not a migration result. 0.5562 is a test/train ratio, not a reconstruction value. Clustering and deduplication are NOT_EVALUATED, which is a different claim from failure.
  • Bridges are directional, and composition is not free: an A_to_B authorization says nothing about B_to_A, and chaining two bridges produces a new derived transformation whose end-to-end behavior has to be measured directly — preservation scores do not multiply along a chain.
  • The registry holds multiple candidate bridges per pair, selects by matching operation, scope, and direction, returns the evidence trace alongside the decision, and defaults to denial whenever no matching active authorization exists. A bridge’s evidence stays pinned to its declared pair and is not falsified by a newer space existing elsewhere. Separately: space incompatibility and encoder obscurity are not privacy boundaries — treat embedding vectors as sensitive data regardless of which encoder produced them.

Next

Chapter 20 decided where evidence belongs in the runtime — which object holds a measurement, which object holds a decision, and how the two are allowed to touch. It did not settle how much any one piece of evidence actually deserves to say. retrieval_ratio, calibration_transfer, relation_profile_corr, and the rest have each been given a precise definition here and in Chapters 18–19, but a precise definition is not the same as a full audit of what a metric can and cannot detect. The next chapter runs that audit directly: reconstruction, neighborhoods, top-k behavior, ranking, task metrics, calibration, relation preservation, hard negatives, round trips — computed, compared, and interrogated one at a time for exactly what each does and does not certify, and what can still go wrong underneath a number that looks reassuring.