AIBussin Book Published

Embeddings From First Principles

Explore what it means for information to become geometry, when that geometry can be trusted, and how representation, similarity, retrieval, calibration, cross-space alignment, and compression combine into an embedding runtime that knows its own limits.

A vector is a list of numbers.

Somewhere between that list of numbers and a working search system, a retrieval-augmented model, or a semantic memory, we start talking about meaning. We say two documents are “close.” We say a query “matches.” We say a space “understands” a distinction.

This book is about that gap.

Not the API call that turns text into a vector — that part is easy. The gap is everything we assume once the vector exists: that distance tracks meaning, that similarity implies equivalence, that a high score implies a correct answer, that two models with the same output dimension produce comparable coordinates, that a smaller representation preserves a larger one.

Each of those assumptions is sometimes true, under conditions, and the conditions are measurable.

What does it actually mean for information to become geometry — and when can we trust that geometry?

That is the question the book works through, one mechanism at a time.


The problem we will explore

The simple picture of an embedding system looks like this:

    graph LR
    T[text] --> E[embedding model] --> V[vector] --> S[similarity] --> R[result]
  

It is a useful starting point and a poor foundation.

In practice, each arrow hides a decision, and each decision has a failure mode:

which transformation produced the vector, under which objective

what geometry that transformation actually imposes on the corpus

how many degrees of freedom the representation really uses

whether the metric matches the distinction you care about

whether a near neighbor is a correct neighbor

whether the operating threshold transfers to new data

whether a second model's vectors mean anything next to the first model's

whether a compressed or translated vector preserved what mattered

These are not the same question, and no single similarity score answers them.

The book therefore expands the picture into a process:

    graph TD
    D[define the relation that matters] --> EM[embed: record the space]
    EM --> G[measure the geometry]
    G --> ST[stress it: hard cases, perturbations, model swaps]
    ST --> C[calibrate: distributions and operating points]
    C --> P[preserve: what survives compression and translation]
    P --> GV[govern: turn measurement into policy]
  

The shift is from asking “are these two things similar?” to asking “under this representation and this metric, what does their geometric relationship establish, and what does it not?”


One corpus, built on for the whole book

The book runs one continuous experiment rather than disconnected notebooks.

We start with a designed set of about 1,200 short text items — call it the RELATE corpus — where the relationship between pairs is known and labeled:

equivalent            paraphrase           entailment
partial-support       contradiction        negation
relation-swap         temporal-mismatch    topic-related
entity-related        unrelated

hard-negative is a role tag for retrieval candidates with an underlying relation, not a base relation label.

The RELATE corpus and its longer-document companion RELATE-DOC are the datasets this book runs on; “RELATE” also names the wider project they belong to, whose relation-specific readout mechanism (Relation Projection) is brought in as external evidence in Chapter 15. Chapter 1 draws the distinction once.

Every few chapters, we learn something new about the same corpus:

Ch 1    cosine looks excellent — on the easy relations
Ch 6    local geometry is stranger than the global picture suggests
Ch 7    a 768-dimensional vector does not use 768 dimensions
Ch 10   the nearest neighbor is confidently, geometrically wrong
Ch 11   hard negatives collapse the margin that easy benchmarks show
Ch 14   the threshold that worked yesterday does not transfer
Ch 15   one scalar cannot carry what the geometry contains
Ch 16   a second model builds a different universe from the same text
Ch 18   we try to translate one universe into the other
Ch 21   the translation preserves retrieval and destroys something else
Ch 22   a bridge retrieves the paired target almost perfectly while reproducing only part of the target neighborhood
Ch 23   preserving source-space cosine trades directly against target fidelity — so a translation must name which geometry it is trying to preserve
Ch 24   we compress a document and measure what the vector forgot
Ch 25   a semantic edit is an operator with a complexity, not a magic direction
Ch 26   we build a runtime that carries all of these limits as metadata

Chapters 22–23 add a second, larger cross-space benchmark — ~2,900 sentences from the frozen corpora, a deterministic train/test split, a 1024-dimensional encoder translated into a much larger decoder-LM embedding space — to sharpen the alignment results: counterpart recovery, target-neighborhood fidelity, and the geometry named by the training objective are measured separately.

That discovery arc is the structure of the book.


The recurring ideas

The book is a sequence of distinctions that are easy to collapse and expensive to collapse:

A vector is not meaning. It is a representation produced by a learned transformation under a particular objective.

Similarity is not equivalence. “Close” is a property of the representation and the metric, not of the two texts.

Proximity is not truth. A false statement can sit next to its correction.

Retrieval is not verification. Finding a passage is not confirming a claim.

Dimension is not information capacity. Nominal dimension and effective dimension are different numbers.

Equal dimensions do not imply compatible spaces. Two 768-dimensional encoders do not share a coordinate system.

A bridge is not compatibility until preservation is measured. A map between spaces can keep rankings while losing calibration, or keep clusters while losing rank order.

Counterpart recovery is not geometry preservation. A bridge can place the true target in the top ten every time while replacing a fifth of the target model’s native neighbors, reordering a third of the survivors, and keeping only half the cluster structure. “Found the right item” and “recreated the target space” are different claims.

“Preserve the geometry” is incomplete until the reference is named. Source-isometry, target-neighborhood fidelity, and downstream task preservation can disagree. A loss can faithfully preserve the wrong geometry.

Identity is exact; compatibility is empirical; usability is a policy decision. These are three separate layers:

SPACE IDENTITY   exact, configuration-derived      → space_hash
COMPATIBILITY    empirical, task-dependent          → measured preservation / evaluation
USABILITY        a policy decision                  → usable_for(scope, operating_point)

A matching space_hash establishes declared identity, nothing more. Whether two spaces are compatible for a task is measured, not inferred from the hash. Whether a system may use that compatibility is a scoped policy call.

Every transformation of an embedding creates an obligation to measure what was preserved. Truncation, whitening, a cross-space bridge, document compression, a semantic edit — each returns different vectors, and none is equivalent to its input until a preservation profile says so.

Coarse geometric preservation is systematically compatible with fine semantic failure. The book’s experiments show this three independent ways: an encoder places a sentence and its negation in nearly the same spot (Ch 1, 10); a cross-space bridge preserves retrieval to within a few points while inverting the paraphrase-vs-negation distinction (Ch 21 — measured, +0.033 → −0.107); and a document compression holds its position in semantic space to three decimal places while silently reversing which company acquired which (Ch 24 — whole-document embedding drift detected 0% of every controlled corruption). A representation can look preserved and be wrong about the one thing that matters. Only a claim-level check finds it.

These distinctions converge on one idea, which the book earns rather than asserts:

Geometry is evidence about a representation, not permission to use it.

They are stated early and paid off with experiments.


What the book is designed to teach

By working through the chapters, you should be able to:

  • distinguish identifiers, features, vectors, representations, and embeddings, and say what a learned objective does and does not put into a vector;
  • build a small embedding space by hand and watch relationships emerge from co-occurrence and prediction;
  • derive cosine, dot product, and Euclidean distance and explain why normalization changes the answer;
  • measure intrinsic dimension, effective rank, participation ratio, and anisotropy, and compare the geometry two models impose on the same corpus rather than their advertised dimension;
  • implement retrieval from first principles and characterize how top-k, thresholds, and approximate search change the resulting “memory”;
  • construct paraphrases, negations, contradictions, and hard negatives, and measure where naive similarity ranks them;
  • evaluate an embedding with Recall@k, MRR, and nDCG while separating representation quality from application quality;
  • build positive and negative score distributions, choose false-acceptance-calibrated operating points, and identify ambiguity bands;
  • represent a pair with several geometric signals — margin, local density, neighborhood stability — instead of one similarity number;
  • define a space identity (model, version, dimension, normalization, configuration, hash) and reason about coexistence and re-embedding when a model is upgraded;
  • learn linear and orthogonal maps between two embedding spaces, and distinguish coordinate reconstruction from semantic preservation;
  • specify an embedding bridge with explicit usable_for scopes and preservation metrics rather than an assumed universal compatibility;
  • distinguish paired-target recovery from target-neighborhood preservation, and evaluate the two separately;
  • choose whether a translation should preserve source structure, imitate target geometry, or optimize a downstream task relation — and state that choice explicitly;
  • measure whether a compressed representation preserved the geometry of the original document; and
  • assemble an embedding runtime that answers “what space produced this vector, is it compatible with that one, how stable are its neighbors, and what does a translation preserve?”

The objective is not a magic similarity score. It is to understand the representation layer well enough to know what a number means, where it breaks, and what a system should do with it.


Where this book sits

The From First Principles books share a pattern. Each one refuses a comfortable shortcut:

Models          a model is a learned function, not a knower of facts
Hallucination   a model's output is not automatically evidence
Agents          a model call is not automatically an agent
Context         more information is not automatically better context
Embeddings      a vector is not meaning

This book adds the representation layer underneath all of them. Retrieval systems, agent memory, RAG pipelines, deduplication, clustering, recommendation, and semantic caching all run on embeddings, and all inherit the geometry’s limits whether or not anyone measured them.

The book deliberately spends little time on vector databases. FAISS, Qdrant, and pgvector are implementations of one primitive operation. The subject here is larger and outlives them: representation, geometry, measurement, retrieval, failure, calibration, alignment, translation, compression, and infrastructure.


The promise

By the end of Embeddings From First Principles, you should not see an embedding as a black-box array that “captures meaning.”

You should see a conditional, lossy, model-specific transformation whose geometry can be measured, stress-tested, calibrated, versioned, translated, and compressed — and whose trustworthiness for a given job is a number you can compute rather than a hope you carry. And you should treat every operation that transforms that geometry — a truncation, a bridge, a compression, a semantic edit — as something that owes you a measurement before you trust its output.

The book’s capstone is a runtime that carries those limits as metadata. Before it, the close of Part VI sharpens two questions the runtime must ultimately answer: did a translation merely recover the paired item or recreate the destination structure (Chapter 22), and what geometry was the translator actually trained to preserve (Chapter 23)?

The book begins with three words and a list of numbers.

It ends with a system — and a reader — that can ask not only whether a transformed representation looks close, but which property survived, against which reference, and whether that is the property the destination actually needs.

Contents

Chapters

01

What Is an Embedding?

Start from the list of numbers. Separate identifiers, features, vectors, representations, and embeddings, and establish the load-bearing claim of the book: an embedding does not contain meaning — it is a representation produced by a learned transformation under a particular objective.

02

Meaning Becomes Geometry

Build a tiny embedding space by hand and turn semantic questions into geometric ones — distance, direction, angle, magnitude, neighborhood. Then use a measured projection experiment to mark exactly where that translation can be trusted and where it only looks trustworthy.

03

Learning an Embedding Space

Refuse the handout and build a small embedding space from a corpus: within-sentence co-occurrence, PPMI weighting, then a truncated factorization. Watch a usable geometry appear, then inspect the construction closely enough to predict what it destroys before measuring it on RELATE.

04

Similarity Is a Decision

Derive cosine, dot product, and Euclidean and Manhattan distance as invariance choices, prove that L2 normalization makes three of them rank identically, then run a metric sweep on RELATE where the expected effect almost vanishes — and use that negative result to show that how much metric choice matters is itself a property of the representation.

05

Dimensions Do Not Mean What You Think

Ask what dimension 173 means and find that the question is not basis-invariant. Rotate a whole embedding space by an orthogonal matrix: every coordinate changes, and cosine, distance, and retrieval do not. Separate basis dependence from superposition, and anisotropy from both.

06

Neighborhoods and Hubs

Inspect local structure directly — the directed k-NN graph, in-degree, hubs and anti-hubs, local density. Measure hubness on RELATE, predict that removing hubs will help retrieval, watch it slightly hurt instead, and separate the detector from the policy.

07

How Many Dimensions Does Meaning Need?

One 768-dimensional representation, described by numbers from about 4 to 768 depending on what you measure. Separate ambient size, spectral size, local intrinsic-dimension estimates, and task-retention thresholds — and discover that no single whole-space geometry number supplies the task-safe compression dimension.

08

The Shape of an Embedding Space

Freeze the corpus, change only the model, and measure the geometry: common-direction bias, centered-covariance concentration, distance concentration. Then reshape one space by centering and whitening — and find that a near-zero random-pair cosine is not a quality target.

09

From Similarity to Search

Define retrieval as a policy — query representation, candidate eligibility, scoring, ranking, cutoff — before optimizing how it runs. Then measure an HNSW index against the exact reference on RELATE, and separate implementation fidelity from semantic correctness.

10

The Nearest Neighbor Can Be Wrong

A vector score can prefer a candidate that is on-topic, lexically similar, geometrically close, and wrong in polarity, role, time, or claim. Name the ways this happens, then use RELATE typed hard negatives to measure whether cosine — or a second-stage entailment score — separates a known-good candidate from the near miss.

11

Hard Negatives

The same model, the same queries, the same reference positive — only the negative-selection rule changes, and the measured margin moves from 0.47 to 0.06. Define margin precisely, separate selection, geometric, and semantic hardness, and show that the negative set is part of the measurement instrument, not an implementation detail.

12

Retrieval Is a Policy

A retrieval system is not an embedding model. It is a versioned composition: representation, eligible candidates, scoring and fusion, selection, execution strategy, and context assembly. Measure a three-policy ablation on RELATE, discover that its hard-negative metric never observes the reranking stage, and learn the rule that generalizes: an ablation is only valid if the metric passes through the stage you changed.

13

How Do You Evaluate an Embedding?

An nDCG score is not a property of a model. It belongs to a query population, a candidate universe, a relevance definition, a metric, an aggregation rule, and a model-use protocol. Measure the same three encoders on RELATE v0.1 and v0.2 — same item pool, same metric implementation — and watch the aggregate winner change when the query workload expands.

14

Calibration

A cosine of 0.81 is a geometric score, not an operational verdict. Define a task-specific positive and negative class, build their score distributions, and derive an operating point from AUC, an approximate equal-error threshold, and a quantile-based abstention band — then compare separately fitted operating points across domains.

15

Is Similarity One-Dimensional?

One cosine score tells you how a pair aligns. It does not tell you how decisive that alignment is relative to the candidates around it. Separate absolute pair score from relative-ranking and neighborhood evidence, measure what a supervised diagnostic built from that evidence actually buys on RELATE hard negatives, and keep evidence and routing policy in separate objects.

16

Change the Model, Change the Universe

Two encoders can produce vectors of the same length without producing vectors in the same coordinate frame. Learn why equal dimension is syntax, not compatibility, and how to measure structural and decision agreement between independently trained spaces without ever comparing their coordinates directly.

17

Versioning the Space

A vector is not just float[d]; operationally it is float[d] plus the identity of the transformation that produced it. Give that identity an exact, content-derived hash, separate it cleanly from compatibility and usability, and read a deliberately dangerous mixed-index probe that shows why the separation has to be enforced before a quality metric ever gets the chance to reveal the damage.

18

Can One Embedding Space Be Translated Into Another?

Suppose two models represent the same objects. Can we learn a map T such that T(E_A(x)) ≈ E_B(x)? Fit the simplest plausible transformation, test it on entity families it never saw, and discover that ≈ is not one number — it is a preservation contract that different properties satisfy at radically different levels.

19

Alignment

Hold the anchors, the split, and the evaluation fixed. Change only the alignment method — the null map, orthogonal Procrustes, affine ridge, a small MLP — and discover that no method wins every preservation property. Alignment is a toolbox of inductive biases, not a ladder from weak to strong.

20

The Embedding Bridge

A map is a function. A bridge is a map plus exact identity, derivation provenance, held-out preservation evidence, and an operation-specific authorization contract — so a fitted transformation never silently becomes a permission. Package the measured MiniLM-to-mpnet Procrustes evidence into an auditable ALLOW / DENY / CONDITIONAL / NOT_EVALUATED decision, and build the registry that fails closed by default.

21

Did the Bridge Preserve the Space?

A preservation metric measures one property, not bridge quality. Audit the MiniLM-to-mpnet Procrustes bridge's eight measured fields against their actual implementations, discover that a reassuring 0.86 relation-profile correlation hides a paraphrase-negation contrast flipping sign, and learn why an artifact's filename is not evidence — the supervised-bridge-ceiling experiment never leaves source space at all.

22

Retrieval Is Not Geometry

A bridge can recover the exact paired target almost every time while reproducing markedly less of the neighborhood around it. Separate counterpart recovery from structural fidelity on the same translated vectors, attack the interpretation with a high-CKA control and a shuffled-correspondence control, and learn that even a well-defined metric can fail to discriminate what another one discriminates clearly.

23

What Should a Translation Preserve?

A translation objective is incomplete until it names the geometry it is trying to preserve. Hold architecture, data, and seed fixed, sweep the weight on a source pairwise-cosine regularizer, and watch most target-facing Wave 6 metrics decline while cluster ARI refuses to cooperate — a controlled demonstration that stronger pressure toward one invariant can work against properties the destination consumer needs.

24

Can a Smaller Representation Preserve a Larger One?

The text got shorter; the embedding did not. Audit a five-layer preservation stack against its actual implementation and discover that whole-document cosine, self-document recovery, and query-conditioned similarity can all stay silent while a compressed summary reverses a relation or changes a number — and that only an external NLI verifier, itself fallible, catches most of it.

25

From Deltas to Operators

A pairwise delta is one candidate operator, not the definition of a transformation. Fit a complexity ladder — identity through a nonlinear map — hold out by unseen base sentence, and discover a taxonomy that looks clean until you notice it was produced by a 0.85 reconstruction bar: identity can be simplest while delta scores far higher, and NONE can still retrieve its target perfectly.

26

Building an Embedding Runtime

The capstone: RELATE 1.0, the executable embedding runtime this book's research forced into being. Exact space identity, measured geometry, hard negatives, calibration, cross-space comparison, bridges, preservation profiles, compression, and semantic operators — composed through one doctrine: identity is exact, compatibility is measured, usability is policy-scoped, and lineage composes while permission does not.