From Deltas to Operators
A pairwise delta is one candidate operator, not the definition of a transformation. Fit a complexity ladder — identity through a nonlinear map — hold out by unseen base sentence, and discover a taxonomy that looks clean until you notice it was produced by a 0.85 reconstruction bar: identity can be simplest while delta scores far higher, and NONE can still retrieve its target perfectly.
Part VII — What Survives Transformation
The other thing a subtraction might mean
king − man + woman ≈ queen is the famous demonstration that a direction in embedding space can correspond to a semantic relation. It is also, as Chapter 2 noted, partly curated and works best locally — and it has a long list of documented problems: the offset method’s success is entangled with plain cosine-neighborhood structure, so a “the direction transfers” result has to beat the baseline of ignoring the offset and returning the nearest neighbour of the source word (Linzen, 2016); and performance varies wildly by relation type (Rogers, Drozd & Li, 2017).
The moral is not that directions never work. It is that the difference vector was never, by itself, the right object to evaluate. Define it anyway, because it is a starting point:
Δ(x₁, x₂) = E(x₂) − E(x₁)
For a pair like (rough draft sentence, edited sentence), Δ points from “before” to “after.” But a single arrow between two points cannot tell us whether the rule that produced it is reusable. Does the same arrow work on new content? A vector cannot answer that on its own. It has to be tested as a candidate transformation — one member of a larger family of candidates.
So this chapter replaces the question “does Δ transfer?” with a harder one:
What is the simplest operator class that represents this transformation out of sample?
A genuinely production-ready answer would also need a second property this chapter’s own measured evidence does not establish: that the operator leaves alone whatever it is not supposed to touch. Keep that distinction sharp throughout — what Wave 5 measured is whether an operator can reconstruct held-out paired targets under a fixed criterion; what a reusable production operator would additionally need is evidence that it does not distort inputs it should leave alone. The second is a real and important requirement. It is not part of the pass rule below.
The operator ladder
Start from one general form and read several familiar ideas as points along it:
T(x) = W x + b
| Rung | Operator | What it can do | Parameters |
|---|---|---|---|
| 0 | T(x) = x | nothing — the no-change control | 0 |
| 1 | T(x) = x + Δ | a reusable direction (the classic “delta”) | d |
| 2 | T(x) = x + xW₁ + c, rank(W₁)=1 | a rank-1 residual correction | ~3d |
| 3 | T(x) = x + xWᵣ + c, rank(Wᵣ)=r | a low-rank residual operator (rank r) | ~(2r+1)d |
| 4 | T(x) = Wx + b | a full affine map | d² + d |
| 5 | T_x(x) = W_x x + b_x | a locally weighted affine fit, re-fit per query | varies by neighborhood |
| 6 | T(x) = f_θ(x) | a nonlinear operator (small MLP) | ~10⁶ |
| 7 | T(x, c) | a content-conditioned transformation | conditioning-dependent |
Read the ladder as a complexity and inductive-bias hierarchy, not a strict mathematical nesting of concrete fitted procedures. Identity, delta, and a full affine map genuinely do contain one another as special cases — a delta is an affine map with W = I; a full affine map is the general case rungs 1–3 restrict. But the specific fitted implementations at rungs 5 and 6 — one particular locally weighted regression, one particular MLP architecture and training recipe — are concrete methods with their own inductive biases, not automatically literal supersets of every rung beneath them under their own training procedures. This is the same correction Chapter 19 required for the cross-space alignment ladder, applied here one level down: an operator is a set of constraints on what a transformation is permitted to do, not a rung on a guaranteed monotonic ladder of quality.
Each step up buys expressive freedom and pays in parameters — from d to d² + d to ~10⁶ — which is why a higher rung has to earn its place on held-out content rather than merely fit the pairs it was trained on.
Two connections tie this to Part VI without overclaiming their scope:
- A delta is an affine map with
W = I. Rung 1 is a genuine special case of rung 4 — this containment is exact. - Cross-space bridges (Chapters 18–21) and same-space semantic operators can share the same affine form,
Wx + b, and the same preservation discipline. They do not automatically share an identity contract. Wave 3’s ridge bridge fits scikit-learn’sRidgewith its default fitted intercept —b ≠ 0in general — and a cross-space bridge can be rectangular, mappingd_source → d_target; this chapter’s semantic operators act entirely within one 768-dimensionalall-mpnet-base-v2coordinate system. So “a bridge hasb = 0” is not a general fact, and “alignment and semantic transformation are literally the same operation” overstates a real, useful kinship: they share an operator vocabulary and a preservation problem, acting under different identity contracts.
Where reusable operators plausibly exist — and where they plausibly do not
Before running the test, state the priors. Not every transformation should have a reusable operator. The ladder is a hypothesis space, not a promise.
Plausibly reusable (some rung might clear a bar):
- Editorial transformations. “Wordy → concise,” “passive → active,” “informal → formal.” If a writing system remembers
(source, target)edit pairs, an operator fitted to those could plausibly apply the same edit to new sentences. - Grammatical transformations. “Singular → plural,” “present → past” — highly regular in form; a candidate for a low rung.
- Relation directions. “Country → its capital,” “company → its CEO” — the classic analogy setting.
Plausibly not (no rung likely clears a bar):
- Content-dependent transformations. “Summarize” depends entirely on what is being summarized; there is no single summarize-operator (Chapter 24’s territory).
- Anything the base model does not represent as accessible structure. If polarity sits near the model’s noise floor under cosine (Chapter 10), “negate this” may have no reusable operator at any rung.
- Abstract multi-step reasoning. “Problem → solution” is not one operator.
This section does not establish that any plausible case actually yields a passing operator. It only sorts candidates so the bake-off below can surprise us in an informative way: a pass where none was expected, or a failure where one was, means something different from a routine confirmation.
Borrowed from model internals — and why it does not carry over
There is a large literature showing that inside a language model’s residual stream, some concepts are approximately linear directions (Park, Choe & Veitch, 2024), that a single steering vector shifts behaviour (Turner et al., 2023; Rimsky et al., 2024), and that others need an affine correction rather than a bare direction, or a low-rank weight edit (Meng et al., 2022; 2023).
That work motivates the operator vocabulary. It is not evidence about an external sentence encoder. A model’s hidden states are shaped by the next-token objective and read through the unembedding; a sentence embedding is a different object, with a different training objective and a different geometry (Chapter 1). This chapter tests the ladder on sentence embeddings directly, importing only vocabulary and hypothesized failure modes from the internals literature — not conclusions.
The distinction matters because it blocks the easiest over-claim available here: “steering vectors work inside the model, therefore deltas must transfer outside it.” They may or may not. Measure, do not import.
Knowledge graphs climbed a parallel ladder
Independently, the knowledge-graph embedding community climbed a related progression, each step motivated by a class of relations the previous rung could not represent:
TransE relation = a translation: h + r ≈ t works for 1-to-1 relations,
breaks on 1-to-N / N-to-1 / N-to-N
TransH project onto a relation-specific hyperplane, then translate
TransR apply a relation-specific projection matrix, then translate (a per-relation affine map)
RotatE relation = a rotation in complex space composes and inverts; captures
symmetry / antisymmetry / inversion patterns
This is a genuinely parallel progression from simple to richer relation operators, not the identical mathematical ladder used above — TransE, TransH, TransR, and RotatE have their own objectives, scoring functions, and representation spaces, distinct from the affine-map family this chapter fits. The transferable lesson is narrower than “these models prove sentence transformations need rotations”: different relation patterns forced knowledge-graph models to adopt richer relation-specific transformations when simpler translation models could not express those patterns. That motivates building an operator ladder at all. It is not evidence about which rung any particular RELATE-DOC transformation needs — this chapter borrows the question and the vocabulary, then runs its own bake-off to find out.
The operator bake-off, defined exactly
For a typed transformation with paired examples, fit every rung on the same training pairs and evaluate all of them the same way.
- Collect
npairs{(a₁, b₁), ..., (aₙ, bₙ)}of one transformation type. - Hold out by base sentence — the unique underlying content each pair is built from — so that no test pair shares its base content with any training pair. This tests whether an operator generalizes to unseen content of the same transformation type; it does not test transfer across template families, entity families, or domains, which the corpus records separately but which this bake-off does not hold out by.
- Fit rungs 0–6 on the training split. (Rung 7, the content-conditioned operator, was not built or measured.)
- On the held-out split, compute reconstruction cosine, top-1 paired-target retrieval, and mean displacement for each rung.
- Compare against a baseline, and find the simplest rung that passes a predeclared contract.
The baseline adapts the Linzen critique to this paired-transformation setting: for each held-out source, find its nearest training source by cosine, then reuse that training pair’s target as the prediction — never touching the fitted operator at all. Score that prediction’s cosine against the true held-out target. A learned operator that does not beat this nearest-training-source transfer baseline by the predeclared margin has not demonstrated an improvement over nearest-example transfer under this contract. No random-delta baseline was fit or persisted in this run; it remains a sensible additional sanity control for a future pass, not something reported here.
The pass contract, exactly as implemented: a rung passes for a given transformation if its held-out reconstruction cosine is at least 0.85, and it beats the nearest-training-source baseline’s reconstruction cosine by more than 0.01. That is the entire rule. It says nothing about collateral effects on content the transformation was never meant to touch — no such measurement exists in this artifact, a limitation the chapter returns to explicitly below.
flowchart TD
P["collect n (source, target) pairs of one transformation type"] --> S["hold out by BASE SENTENCE — unseen content, same transformation type"]
S --> F["fit rungs 0-6 on the training split: identity / delta / rank-1 / low-rank / full affine / local affine / MLP"]
F --> BL["baseline: nearest-training-source -> reuse ITS paired training target"]
BL --> W{"lowest rung with reconstruction cos >= 0.85 AND beating the baseline by > 0.01?"}
W -->|"a rung qualifies"| SP["simplest passing operator = that rung, UNDER THIS CONTRACT"]
W -->|"none qualifies"| NP["NONE — no rung clears this reconstruction contract at this data scale"]
What a reusable operator would additionally need — and why Wave 5 does not measure it
The knowledge-editing literature makes a useful demand that this bake-off’s pass rule does not enforce: an edit’s efficacy (did the target change as intended?) has to be weighed against its specificity (did everything the edit was not meant to touch stay put?), and aggressive editing can cause measurable “representation shattering” elsewhere in the space. Carried into this chapter’s vocabulary:
A genuinely reusable transformation should move what it is supposed to move without unnecessarily disturbing the geometry it was never meant to touch.
Wave 5 reports a mean_displacement field for every rung — ‖T(x_test) − x_test‖, averaged over the held-out test sources. Read this precisely: it is computed on the same held-out sources the transformation is meant to move, not on a separate population of items the transformation should leave alone. It tells us how far the operator moves its intended targets. It tells us nothing about specificity, collateral drift on unrelated content, or neighborhood churn elsewhere in the space — no such held-out “unaffected” population exists anywhere in this artifact.
A correctly designed specificity test would need: (1) a defined population of inputs the transformation should not meaningfully affect; (2) the operator applied to that population; (3) displacement and neighborhood-churn measurements on it; (4) an application-specific collateral budget declared in advance. None of that exists here — it is a real, important, and entirely unmeasured requirement for a production-grade operator, marked explicitly PROPOSED throughout what follows. Note too that the “unaffected population” itself is not universal: a transformation like “make all prose more formal” may legitimately be meant to move everything, while “reverse this specific factual relation” should move almost nothing else — the correct unaffected population comes from the transformation’s own declared contract, echoing Chapter 23’s lesson that authority comes from the consuming use, not from the geometry alone.
Three levels of evidence
reconstruction the operator reconstructs held-out targets of the same type weakest, and what Wave 5 measures
↓
transfer it still works under unseen template / entity / domain shift stronger — PROPOSED, not measured
↓
algebra a genuinely inverse map recovers the source; composed strongest — only a narrow diagnostic
transformations match directly observed composed examples measured (below), not full algebra
An operator that reconstructs but has not been tested under template, entity, or domain shift has cleared the weakest bar this hierarchy names. An operator that also transfers under those shifts would be stronger evidence of reusability — a real next step this chapter’s evidence does not reach. An operator whose reverse map genuinely inverts it, and whose composed applications match directly observed composed examples, would be the strongest evidence that the transformation behaves like structured algebra rather than a local curve fit. Hold this hierarchy in view as a standard to measure toward, not as a description of what the demonstration below actually reaches — it reaches only the first level, plus one narrow diagnostic toward the fourth.
Demonstration: running the bake-off
MEASURED on RELATE-DOC v0.1 Family B, Transformation Wave — artifact
experiments/embeddings-from-first-principles/wave5/artifacts/operator-bakeoff.json.all-mpnet-base-v2; nine typed transformations, 260 total pairs, held out by base sentence.
RELATE-DOC v0.1 Family B is a small, synthetic, templated corpus of short sentences (roughly 12–20 words in the transformation pairs). Each transformation type has only 10–45 total pairs, split into training and a held-out test set of 3–15 examples. This is a genuinely small-data regime for the more expressive rungs — call it that explicitly, not “a realistic data volume” in any general sense, and hold every result below to that scale.
Run the full ladder on all nine transformation types and look first at only the columns that decide the pass contract:
transformation n_pairs/n_test identity delta simplest passing (contract: >=0.85, beats baseline)
active_to_passive 45 / 15 0.9509 0.9537 IDENTITY
present_to_past 35 / 11 0.9609 0.9733 IDENTITY
relation_swap 10 / 3 0.9788 0.9761 IDENTITY
claim_strengthened 15 / 5 0.8884 0.9518 IDENTITY
temporal_shift 25 / 8 0.8833 0.9727 IDENTITY
claim_weakened 15 / 5 0.8033 0.9408 CONSTANT DELTA
formal_to_informal 35 / 11 0.7536 0.8328 NONE
verbose_to_concise 35 / 11 0.7141 0.7979 NONE
statement_to_negation 45 / 15 0.6689 0.7948 NONE
Before reading the pattern, one detail matters more than the labels themselves: the classification “IDENTITY” is not a statement that nothing semantically happened. It is a statement that the unmodified source embedding already clears the 0.85 reconstruction contract against the held-out paired target. Look at claim_strengthened: identity scores 0.8884, which is above 0.85 — so identity is the simplest rung that passes, even though the constant delta scores markedly higher at 0.9518. The same pattern appears for temporal_shift: identity 0.8833 passes the bar, delta reaches 0.9727 — noticeably closer to the target — and identity still wins the “simplest passing” label because the contract asks only “does this rung clear 0.85,” never “which rung scores highest.”
Simplest passing is not best scoring. A rung earns the label by clearing a predeclared bar with the fewest parameters — not by being the closest fit available.
For active_to_passive, present_to_past, and relation_swap, the identity scores (0.9509, 0.9609, 0.9788) sit close enough to their delta counterparts that the distinction matters less — these are cases where the unmodified embedding genuinely already sits near the paired target. relation_swap’s 0.9788 deserves its own callout: swapping which party performed an acquisition — “Helios acquired Pine” into “Pine acquired Helios” — barely moves the sentence embedding at all. Say exactly what that means: the embedding representation barely distinguishes the before-and-after pair under this metric. It is not a claim that the assertion is semantically harmless, and it is a direct echo of the same boundary Chapter 24 measured for relation reversal inside a compressed document.
claim_weakened is the one clean non-identity success: identity scores 0.8033, below the bar, and constant delta reaches 0.9408, clearing it — a single reusable direction that the classic delta intuition predicts, and the only transformation in this set where that intuition pans out under this contract. Rank-1 and low-rank land in the same neighborhood (0.9394, 0.9401) but do not improve on delta enough to matter, and delta is simpler, so delta remains the simplest passing rung.
For formal_to_informal, verbose_to_concise, and statement_to_negation, identity clearly fails (0.7536, 0.7141, 0.6689 — the embedding genuinely moved), and no rung clears 0.85: the best results — constant delta or low-rank, at 0.8328, 0.8007/0.7979, and 0.8045/0.7948 respectively — all fall short. Full affine, local affine, and the MLP score substantially worse than delta on every one of these three, typically in the 0.35–0.65 range, with mean displacement near 1.0 on their held-out test sources — a severe generalization failure at this data scale, not evidence that affine or nonlinear operators are intrinsically unfittable for sentence embeddings in general. With only 10–45 total pairs per transformation, the data-to-capacity ratio is especially unfavorable for a 768 × 768 affine map or a network with hundreds of thousands of learned weights. The observed failure is therefore consistent with a data/capacity mismatch, but Wave 5 did not isolate that explanation from other possibilities.
The taxonomy under this exact contract:
IDENTITY: active_to_passive, present_to_past, relation_swap, claim_strengthened, temporal_shift
CONSTANT DELTA: claim_weakened
NONE: formal_to_informal, verbose_to_concise, statement_to_negation
One more pattern in the numbers deserves its own line, because it corrects a tempting overstatement: rungs above constant delta are not universally inert. For temporal_shift, low-rank reaches 0.9761 against delta’s 0.9727 — a small numerical improvement. For statement_to_negation, low-rank reaches 0.8045 against delta’s 0.7948 — also a small improvement, though still below the pass bar. The precise, defensible result is narrower than “higher rungs never help”:
No rung above constant delta changes the simplest-passing classification for any of the nine transformations. It never rescues a transformation that delta leaves below the contract, and it is never the lowest rung required to clear it — even where it edges reconstruction slightly higher.
NONE means the reconstruction contract was not cleared — not that no operator can exist
This is the demonstration’s sharpest correction to a natural misreading, and it deserves its own section rather than a caveat. Look at the same three NONE transformations through a second measured field — top-1 target retrieval among the held-out targets — that plays no role in the pass rule at all:
reconstruction (delta) top-1 target retrieval (delta)
formal_to_informal 0.8328 1.000
verbose_to_concise 0.7979 1.000
statement_to_negation 0.7948 1.000
For every one of these three transformations, the constant-delta operator identifies the correct held-out paired target as its single nearest neighbor among all held-out targets, every single time — even though none of them clears the 0.85 reconstruction threshold. This is Chapter 22’s counterpart-recovery lesson, one chapter later and one level down: counterpart recovery and pointwise geometric reconstruction are different properties, and an operator can succeed completely at the first while failing the second. NONE, under this chapter’s contract, means precisely this and nothing more:
No tested rung, under this training recipe and this small-data regime, reconstructs the held-out target closely enough to clear the
0.85cosine bar. It does not mean no reusable operator exists for this transformation; it does not mean the operator cannot identify the right target; it does not mean rung 7’s content-conditioned map, a different metric, or more data would also fail — none of those were tested.
Say the causal story no more strongly than the evidence supports, too: for these three transformations, the embedding moves enough that identity fails the bar, and none of the tested rungs closes the remaining gap. Calling this “the edit is real but irregular” reaches for an explanation — irregularity — that Wave 5 never independently measured. State the observation, not a diagnosis of its cause.
What the classification depends on
Every label above is the output of a specific chain: all-mpnet-base-v2 + RELATE-DOC v0.1 Family B + this exact held-out-by-base-sentence split + these seven fitted operator implementations + cosine reconstruction as the metric + a 0.85 bar + a 0.01 baseline margin. Change any link and some labels could plausibly move — claim_strengthened (identity 0.8884) and temporal_shift (identity 0.8833) sit close enough to 0.85 that a modestly different bar could reclassify them, though no such threshold sweep was actually run here, and no relabeled result should be reported as though it had been.
The scores are observations. The taxonomy is a policy derived from those scores under one specific pass contract — the same relationship Chapter 14 established between a calibration score and its operating threshold, and Chapter 20 established between preservation evidence and an authorization decision.
“Grammatical edits are identity” is not a scale-independent fact about embedding geometry in general. The precise, defensible statement is: for these held-out mpnet examples, under this contract, identity already clears the reconstruction bar for active/passive, present/past, relation-swap, claim-strengthening, and temporal-shift transformations. That claim is fully scoped to this corpus, this model, this split, and this threshold — and it is exactly as strong as it needs to be.
A reverse-map diagnostic, not a mathematical inverse
For the three reversible transformation types (active_to_passive, present_to_past, relation_swap), Wave 5 fits a separate backward ridge map, T_BA, from targets back to sources, alongside the forward map T_AB — and evaluates the round trip on a different, random 75/25 split, seed 0, drawn from that transformation’s own pairs, not the base-sentence-held-out split used everywhere else above. T_BA is not the matrix inverse of T_AB; it is an independently fitted model with its own training data. Write it as T_BA(T_AB(x)), never as T⁻¹(T(x)), to keep that fact visible in the notation itself.
transformation forward T_AB(x) vs native target T_BA(T_AB(x)) vs native source
active_to_passive 0.3626 0.3021
present_to_past 0.2944 0.2429
relation_swap 0.2212 0.2005
Both columns are low under this small-sample run — worth reporting exactly as that, without converting the low numbers into a “percentage of signal retained” reading, which cosine does not support. And the two columns compare against different reference endpoints: the forward column compares against the native target vector; the round-trip column compares against the native source vector the process started from. This is the same lesson Chapter 21 taught about forward-hop cosine and round-trip cosine in a cross-space bridge — do not subtract these two numbers as though the difference measured “additional loss.” The safe, narrow conclusion:
Under this small-sample diagnostic, on a different random split from the main taxonomy, both the forward ridge fit and the separately fitted forward-plus-reverse round trip score low. This says something about these two independently fitted maps on this data; it is not a measurement of a true mathematical inverse, and it does not, by itself, diagnose why the round trip performs the way it does.
Composition was not measured. No result in this artifact chains two transformations — for instance, present_to_past followed by active_to_passive — and compares the result against directly observed doubly-transformed examples. The code that produced this artifact discusses the idea in a comment; nothing persisted evaluates it. Treat composition consistency as an open, unmeasured question, not a finding — and do not read the low forward/round-trip numbers above as evidence that composition specifically fails, since composition and a reverse-map round trip are different operations.
Operators as retrieval keys — a design idea, not a result
For transformations with a genuinely non-identity, non-trivial passing operator, an editorial memory could in principle index and retrieve stored edits by their fitted operator, not just by textual similarity to the source:
new edit request: (source sentence s, desired transformation r)
candidate past edits = retrieve stored (a, b) examples where
their transformation/operator observation is close to r (same KIND of edit)
OR cos(E(a), E(s)) is high (similar CONTENT)
-> operator-based retrieval would surface same-kind edits;
source-based retrieval surfaces similar-content edits;
they are complementary in principle
The measurable claim behind this design would be: for a relation with a genuinely passing, non-trivial operator, operator-similarity retrieves same-kind edits better than source-similarity alone. This was not tested on RELATE-DOC v0.1, and it could not meaningfully be tested here: this bake-off found exactly one non-identity passing operator — claim_weakened’s constant delta — leaving no second or third non-trivial operator to compare it against, and no basis for a comparative retrieval experiment. This is not evidence that operator-similarity retrieval doesn’t work; it is evidence that this corpus, under this contract, never produced the multi-operator situation the idea needs to be tested at all. Record the idea as a design principle worth testing on a corpus or contract that yields more than one passing operator — operator_retrieval_evidence: NOT_MEASURED — rather than enabling it by default anywhere downstream.
What this chapter establishes and what it does not
Establishes: a pairwise delta is one candidate operator among several, evaluated by fitting a complexity ladder and testing reconstruction on unseen content; the exact measured ladder — identity, constant delta, rank-1 residual affine, rank-8 residual affine, full affine ridge, locally weighted affine, and a one-hidden-layer MLP with two Linear layers, with rung 7 (content-conditioned) unbuilt; the exact held-out protocol (by base sentence, not template, entity, or domain); the exact pass contract (reconstruction cosine ≥ 0.85 and beats a nearest-training-source baseline by > 0.01), with no collateral or specificity requirement inside it; on this benchmark, the classification is five IDENTITY, one CONSTANT DELTA, and three NONE, and every label is contract-dependent — identity’s “simplest passing” status coexists with delta scoring substantially higher on two of the five IDENTITY rows; NONE coexists with perfect top-1 target retrieval on all three NONE rows, showing counterpart recovery and reconstruction fidelity are different properties; no rung above constant delta changes any classification, though low-rank sometimes nudges reconstruction slightly higher without rescuing a failing transformation; full affine, local affine, and the MLP generalize far worse than delta on every one of the nine transformations at this 10–45-pair data scale; a reverse-map round-trip diagnostic on three reversible transformations, under a different random split, scores low on both its forward and round-trip legs, which reference different endpoints and should not be subtracted; and operator-key retrieval, and true composition consistency, were both design ideas that Wave 5’s data could not put to a meaningful test.
Does not establish: that any transformation has a reusable mid-complexity geometric operator at any larger data scale — only that none did at 10–45 pairs here; that expressive operators are intrinsically unfittable for sentence-embedding transformations in general; that the taxonomy is threshold-independent, model-independent, or corpus-independent; that a true mathematical inverse or genuine algebraic composition was measured; that collateral specificity on unaffected content was measured, at all; that template, entity, or domain transfer was measured; that operator-similarity retrieval works or fails; or any statistical uncertainty, repeated-seed, or repeated-split result — none exists in this artifact.
Lab 25: reproduce the exact bake-off, and its exact boundary
MEASURED — artifact
experiments/embeddings-from-first-principles/wave5/artifacts/operator-bakeoff.json(Family B,all-mpnet-base-v2, rungs 0–6, held out by base sentence). REPRODUCIBLE —python run_wave5.py.
Question. For each typed transformation, what is the simplest rung that reconstructs held-out targets under a fixed contract — and what does that contract’s exact wording leave unmeasured?
Step 1 — freeze provenance. RELATE-DOC v0.1 Family B, 260 total transformation pairs across nine typed transformations, all-mpnet-base-v2, 10–45 pairs per type. State plainly: a small-data regime for every rung above rank-1.
Step 2 — define the actual held-out test. By base sentence: no test pair shares its underlying content with any training pair, within the same transformation type. This tests content generalization within one transformation type — not template transfer, entity transfer, or domain transfer, all of which the corpus records but this split does not hold out by.
Step 3 — define the exact pass contract. reconstruction_cos >= 0.85 and reconstruction_cos > ignore_baseline_recon + 0.01. State explicitly: top-1 target retrieval and mean displacement are reported for every rung but play no role in pass/fail.
Step 4 — define the baseline exactly. For each held-out source, find its nearest training source by cosine; reuse that pair’s training target as the prediction; score against the true held-out target. This is a nearest-training-source transfer baseline, not “return the source’s own nearest neighbor.”
Step 5 — fit the measured ladder: identity, constant delta, rank-1 residual affine (ridge-then-SVD-truncated, λ=10), rank-8 residual affine (same construction, rank 8), full affine (Ridge(alpha=10.0)), locally weighted affine (k=24 nearest training sources, cosine-weighted), and a one-hidden-layer GELU MLP with two Linear layers (768 → 512 → 768, 400 epochs, Adam, lr=1e-3, weight_decay=1e-5, seed 0). Rung 7 was not built.
Step 6 — reproduce the full nine-transformation table, all nine rows, as in the Demonstration above, including n_pairs/n_test for each row.
Step 7 — interpret IDENTITY correctly. claim_strengthened (identity 0.8884, delta 0.9518) and temporal_shift (identity 0.8833, delta 0.9727) both pass at identity despite delta scoring meaningfully higher. State plainly: identity is simplest-passing here, not best-scoring.
Step 8 — interpret CONSTANT DELTA correctly. claim_weakened: identity 0.8033 fails, delta 0.9408 passes and is the lowest rung to do so.
Step 9 — interpret NONE correctly, using the second metric. formal_to_informal, verbose_to_concise, statement_to_negation all fail 0.85 reconstruction under every rung — yet delta’s top-1 target retrieval is 1.000 for all three. State the conclusion precisely: NONE is a reconstruction-contract failure, not evidence against operator existence, and not a failure of counterpart identification.
Step 10 — disclose the data limits without softening them. n_test ranges from 3 (relation_swap) to 15 (active_to_passive, statement_to_negation). No confidence interval, bootstrap, or repeated split exists for any reported value.
Step 11 — inspect displacement honestly. Report mean_displacement as displacement on the targeted held-out test sources — never as collateral drift, which would require a held-out population the transformation was never meant to move.
Step 12 — reproduce the reverse-map diagnostic, noting its different random split and its two non-comparable reference endpoints, exactly as in the Demonstration above.
Step 13 — mark composition explicitly unmeasured. No composed-transformation result exists in this artifact.
Step 14 (PROPOSED — no artifact backs this) — a specificity test. Define an unaffected population per transformation, apply the fitted operator, and measure displacement and neighborhood churn against a predeclared collateral budget.
Step 15 (PROPOSED — no artifact backs this) — broader transfer. Repeat the bake-off holding out by template family, by entity family, and by domain, using the splits RELATE-DOC v0.1 already records but this bake-off did not use.
Step 16 (PROPOSED — no artifact backs this) — robustness. Repeated MLP seeds; repeated base-sentence partitions; a data-scale curve (hundreds of pairs per type instead of 10–45); a sensitivity sweep on the 0.85 pass bar. None of these results currently exist.
Try it yourself
Add one transformation type from your own domain. Hold out by underlying content — not random rows — so a test example never shares its source content with a training example. Fit identity, delta, and at least one richer candidate. Predeclare your reconstruction metric and pass bar before looking at the results, and include the nearest-training-source baseline. Report every rung’s score, even the ones that fail — a clean
NONE, reported with its exact contract and baseline, is a real result. Then add a second metric, such as top-1 target retrieval, and check whether your pass/fail label would change under a different, equally reasonable evaluation contract. If you want to go further: define a population your transformation should not affect, and measure how far the fitted operator moves it — that is the specificity test this chapter’s own evidence never ran.
Companion component: the transformation observation
This artifact records what was measured, under exactly which contract — never an authorization decision. It stays structurally parallel to Chapter 24’s compression preservation observation and Chapter 20’s evidence/authorization split.
transformation_observation:
transformation_id:
identities:
input_space_hash: # the exact declared source embedding identity, Ch17
operator_parameter_hash: # the fitted operator's own parameters, hashed
derived_output_space_hash: # T(x) is a derived representation, never native identity by default
relation: # e.g. claim_weakened, statement_to_negation
data:
corpus_hash:
n_pairs:
n_test:
split:
type: held_out_by_base_sentence
note: "content generalization within one transformation type only —
NOT template / entity / domain transfer"
evaluation_contract:
reconstruction_metric: cosine_to_paired_target
reconstruction_bar: 0.85
baseline: nearest_training_source_then_paired_target
baseline_margin: 0.01
contract_version:
candidates:
identity: { reconstruction_cos, top1_target_retrieval, mean_displacement_on_targeted_sources }
constant_delta: { ... }
rank1_affine: { ... }
lowrank_affine: { ... }
full_affine: { ... }
local_affine: { ... }
mlp: { ... }
simplest_passing_operator_under_contract: # NOT "intrinsic transformation complexity"
reverse_map_diagnostic_ref: # separate split, separate reference endpoints — see note above
evidence_gaps:
unaffected_item_specificity: NOT_MEASURED
template_transfer: NOT_MEASURED
entity_transfer: NOT_MEASURED
domain_transfer: NOT_MEASURED
composition: NOT_MEASURED
operator_retrieval: NOT_MEASURED
sample_efficiency: NOT_MEASURED
seed_stability: NOT_MEASURED
uncertainty:
repeated_splits: false
repeated_mlp_seeds: false
bootstrap: false
provenance:
artifact_ref:
code_hash:
Two design choices matter here. First, simplest_passing_operator_under_contract is named the way it is deliberately — never “transformation complexity” as an intrinsic property, always a label produced by a specific contract that travels with it. Second, evidence_gaps makes absence first-class and queryable, the same discipline Chapter 20 applied to NOT_EVALUATED: a downstream system asking “does this transformation have a specificity guarantee?” gets an honest NOT_MEASURED, not silence and not a false negative.
Applying an operator creates derived-representation provenance, exactly as a cross-space bridge does (Chapters 17–18): T(x) carries derived_output_space_hash, distinct from the input’s own identity, even when source and output remain in the same nominal 768-dimensional coordinate system. And this evidence is bound to the exact identity pair it was fit and measured on — a checkpoint change, a normalization change, a prefix change, or any other shift in the embedding pipeline does not falsify a transformation_observation’s recorded evidence; it simply means that evidence no longer applies to the new identity, and a new observation is needed before the operator can be trusted there.
Authorization stays a genuinely separate object, referencing this observation rather than being derived from it automatically:
operator_authorization:
consumer:
operation:
requirement_ref:
transformation_observation_ref:
decision: ALLOW | CONDITIONAL | DENY | NOT_EVALUATED
Nothing in the empirical record above says usable_for or not_usable_for. Whether a NONE result blocks a specific downstream use, whether a passing CONSTANT DELTA operator is cleared for batch rewriting, and whether specificity evidence is required before any of that — those are Chapter 20’s questions, asked against this chapter’s evidence, never answered inside it.
Failure modes
- Treating the delta as the transformation. It is rung 1 of a ladder; fit the ladder before trusting any single arrow.
- Treating the ladder as one strict mathematical nesting. Identity, delta, and full affine genuinely nest; one particular local-regression or MLP implementation is an inductive bias, not a guaranteed superset.
- Claiming every cross-space bridge has
b = 0. Wave 3’s ridge bridge has a fitted intercept; only some bridge constructions are origin-constrained. - Claiming Wave 5 held out by template, entity, or domain. It held out by base sentence only.
- Reporting a random-delta baseline. None was fit or persisted; only identity and the nearest-training-source baseline were measured.
- Misdescribing the ignore baseline as “the source’s own nearest neighbor.” It reuses that neighbor’s paired training target, not the neighbor itself.
- Adding a collateral or specificity term to the measured pass rule. The rule is reconstruction ≥ 0.85 and beats the baseline by > 0.01 — nothing else.
- Calling
IDENTITY“nothing changed.” It means the unmodified embedding already clears this contract — not that the edit had no semantic effect. - Calling
NONE“no reusable operator can exist.” It means no tested rung, at this data scale, cleared this reconstruction contract — and all threeNONErows retrieve their correct paired target perfectly under the constant-deltatop1_target_retrievalmeasurement. - Ignoring the metric disagreement inside the
NONErows. Reconstruction failure and counterpart-recovery success coexist there — report both. - Calling
mean_displacementcollateral damage. It is measured on the targeted held-out sources, not a separate unaffected population. - Treating the classification as threshold- or model-independent. It is a direct function of the
0.85bar, the0.01margin,all-mpnet-base-v2, and this exact corpus. - Claiming higher rungs never numerically improve on delta. Low-rank nudges reconstruction upward on some rows; it simply never changes which rung is simplest-passing.
- Generalizing “expressive operators are unfittable” beyond this data scale. Only 10–45 pairs were tested; the result is consistent with a data/capacity mismatch, not a theorem about affine or nonlinear maps.
- Writing the reverse-map diagnostic as
T⁻¹(T(x)). It is a separately fitted backward model,T_BA(T_AB(x)), evaluated on a different random split. - Subtracting forward cosine from round-trip cosine. They reference different endpoints — the native target versus the native source.
- Claiming composition was measured. No composed-transformation result exists in this artifact.
- Calling operator-key retrieval a tested result. It was never run; the corpus never produced a second non-trivial passing operator to test it against.
- Storing
usable_forinside the transformation observation. Authorization is a separate, Chapter 20 object. - Applying a fitted operator under a different embedding identity without new evidence. The old observation’s evidence remains true for its original identity pair; it does not transfer automatically.
What this chapter established
- A pairwise delta is one candidate operator, evaluated against an explicit complexity ladder from identity through a nonlinear map — a hierarchy of inductive biases, not a guaranteed mathematical nesting of every concrete fitted method.
- The measured contract is exact and narrow: held out by base sentence, reconstruction cosine ≥
0.85, beating a nearest-training-source baseline by >0.01— with no collateral or specificity term in the rule at all. - Every label in the taxonomy is a function of that contract, not a property of the edit.
claim_strengthenedandtemporal_shiftpass at identity despite delta scoring far higher, and the threeNONEtransformations retrieve their paired target perfectly (top-11.000) while failing reconstruction at every tested rung.NONEmeans no tested rung cleared this bar — never that no operator exists. - No rung above constant delta ever changed a classification, and the high-capacity fits generalize far worse at this small-data scale — a result about this data regime and this recipe, not a verdict on expressive operators.
- The reverse-map diagnostic is not a true inverse, its two legs reference different endpoints and cannot be subtracted, and composition and operator-key retrieval were design ideas the data could not meaningfully test. The transformation observation records its contract and its evidence gaps as first-class fields, distinct from a Chapter-20 authorization.
Next
Every part of this book has produced an artifact — space records, calibration records, evaluation observations, bridges, preservation profiles, compression observations, and now this transformation observation, each carrying its own contract, its own evidence, and its own explicit gaps. The operator bake-off leaves us with something more useful than a table of deltas: a record of exactly what was tried, what passed, what failed, and under which contract — with NONE stored as first-class evidence rather than silence. The remaining problem is not inventing another metric. It is making a running system refuse to use a representation when the required identity, evidence, lineage, or authorization is missing — turning measurement from advice a careful engineer might consult into a prerequisite the runtime itself enforces. That is Chapter 26.