← Embeddings From First Principles

What Is an Embedding?

Start from the list of numbers. Separate identifiers, features, vectors, representations, and embeddings, and establish the load-bearing claim of the book: an embedding does not contain meaning — it is a representation produced by a learned transformation under a particular objective.

Part I — A Vector Is Not Meaning

Three words and a list of numbers

Take three words:

cat
dog
airplane

Turn them into numbers. Any embedding API will do it. You get three arrays, each maybe 384 or 768 or 1,536 floats long:

cat       [ 0.021, -0.114,  0.062, ... ]
dog       [ 0.019, -0.098,  0.071, ... ]
airplane  [-0.087,  0.203, -0.041, ... ]

Compute the angle between cat and dog. It is small. Compute the angle between cat and airplane. It is larger. Something about “cats and dogs are both pets” appears to have survived the trip into number-space.

It is tempting to stop here and say: the embedding captured the meaning of the words.

That sentence is the first thing this book takes apart — not because it is flatly wrong, but because it hides every decision that produced the numbers, and those decisions determine whether the next distance you compute is useful evidence or just an attractive number.

What is actually in that array of numbers, and what put it there?

Answering that properly takes the rest of this chapter. By the end, the opening example will look less like “meaning turned into numbers” and more like something precise: a learned transformation preserved one kind of relationship strongly enough for geometry to expose it.

The vocabulary we will keep separate

Six terms get used as if they were interchangeable. The book keeps them apart because the differences become load-bearing in later chapters.

  • Identifier. An arbitrary label. Token ID 4021 for cat. It supports equality and nothing else — 4021 is not “closer to” 4022 in any meaningful way. One-hot vectors are identifiers wearing a vector costume: every pair is equidistant.
  • Feature. A measured, named property. has_fur = 1, can_fly = 0, word_length = 3. Features are interpretable by construction; a human decided what each slot means.
  • Vector. Any element of a vector space: a list of numbers you can add, scale, and take dot products of. A vector carries geometric structure but no inherent semantics.
  • Representation. A vector produced from an input by some transformation, chosen so that geometric operations on the vector stand in for operations on the input. The transformation is the point.
  • Embedding. A representation that maps objects into a continuous vector space, learned so that a chosen relationship becomes geometrically useful structure — usually proximity. An embedding need not reduce dimensionality; a one-hot input mapped to a dense vector of the same length is still an embedding. What matters is that the space is continuous, the mapping is learned, and the geometry is meant to be used.
  • Latent representation. An intermediate activation inside a larger model — a hidden layer’s output — that can be used as an embedding. Not every hidden activation is one in the operational sense this book uses: it becomes an embedding only when we commit to reading its geometry as a stand-in for relationships between inputs.
    flowchart LR
    ID["identifier — arbitrary label; supports equality only (token ID 4021)"] --> FE["feature — measured, named property (has_fur=1); interpretable by construction"]
    FE --> VE["vector — element of a vector space; add / scale / dot product; geometry, no semantics"]
    VE --> RE["representation — a vector from an input via a transformation, so geometry stands in for input operations"]
    RE --> EM["embedding — a learned representation into a continuous space where a chosen relationship becomes usable geometry"]
    EM --> LR["latent representation — an intermediate activation read as an embedding once we commit to its geometry"]
  

The move from identifier to embedding is the move from “these are different” to “these differ in graded, structured ways.”

Keep this distinction. An identifier answers which one: cat might be token 4021, but 4021 is not semantically nearer to 4022 than to 88,000. An embedding answers how this item sits among others: perhaps cat lies near dog, farther from airplane, and somewhere between tiger and pet. Equality is the identifier’s main semantically legitimate operation. An embedding gives us a geometry — and the rest of the book is about what that geometry does and does not justify.

An embedding does not contain meaning

Here is the claim the rest of the book leans on.

An embedding does not contain meaning. It is a representation produced by a learned transformation under a particular objective.

Three consequences follow immediately.

The objective decides what is preserved. A model trained to predict neighboring words builds a space where distributionally similar words are close — so good and bad end up near each other, because they appear in nearly identical contexts. A model trained on question–answer pairs builds a space where a question is close to its answer, which are distributionally dissimilar. Same input text, different objective, different geometry, opposite notion of “similar.”

The transformation is selective, and operationally lossy. It is tempting to argue that squeezing a 50,000-word vocabulary, or the space of all English sentences, into 768 numbers must lose information by sheer counting. That argument is too loose: even a low-dimensional continuous space can, in principle, give every item in a finite vocabulary a distinct coordinate. The useful question is not whether the mapping can keep every item distinct. It is which distinctions the learned geometry makes accessible to the operations we intend to use. Training rewards some relationships more directly than others; distinctions it does not reward may collapse, become weak, or survive only in directions or readouts that cosine similarity does not expose. In that operational sense, the representation is lossy: it does not preserve every distinction in a form the downstream geometric operation can reliably recover.

The geometry is conditional. “cat is near dog” is a fact about this model’s output space under this metric. Change the model, the metric, or the normalization and the statement can change.

None of this makes embeddings less useful. It makes them a tool with a spec sheet instead of a magic trick.

This is the first form of an idea the book keeps sharpening: geometry is evidence about a representation, not permission to use it. A small angle between two vectors is a fact about where a particular transformation placed them. Whether that fact licenses a decision — “these are duplicates,” “this passage answers the query,” “this translated vector is as good as a native one” — is a separate question, and the answer is a measurement, not an assumption.

Rule of thumb for the whole book. Before you interpret a distance, identify what produced the space: the model, what you know about its training signal, the normalization, and the metric. If some of that is unknown, treat the geometry as an empirical object: measure what it does on the relationship you actually care about.

Demonstration: the RELATE corpus

MEASURED on RELATE v0.1 (corpus_hash 8cad6816…eda525b3), Wave 1 row 1.1 — artifact experiments/embeddings-from-first-principles/wave1/artifacts/relation-cosine-by-type.json. Five sentence encoders; the table shows bge-large-en-v1.5.

The last section was an argument. This one is a measurement. If the objective really does choose what the geometry makes easy to read, then we should be able to construct pairs that are nearly identical in wording but different in the relationship a careful reader cares about — and watch the geometry struggle.

Throughout the book we work with one designed dataset, the RELATE corpus: 1,173 short text items with 1,181 pairs, each labeled by the relationship between the two sides —

same or overlapping claim:          equivalent   paraphrase   entailment   partial-support
incompatible or changed assertion:  contradiction   negation   relation-swap   temporal-mismatch
related without claim equivalence:  topic-related   entity-related
no relevant relation:               unrelated

Those are the eleven relation labels used here. hard-negative is different: it is a role tag for a candidate that looks plausible under lexical, topical, or entity overlap but is not a correct match. A hard negative still carries an underlying relation such as negation, relation-swap, or temporal-mismatch. That distinction matters because those are exactly the cases that will expose what generic similarity misses.

A note on the name. RELATE is an umbrella research project on relation-specific readouts of embeddings; the corpus spec expands it as Relations Explicitly Labeled And Typed for Embeddings. Within that project, the RELATE corpus and its longer-document companion RELATE-DOC are the evaluation datasets, and they are what this book uses from here on. The project’s other half — Relation Projection, a mechanism that trains a readout for one chosen relation over a frozen embedding — is a separate line of work with its own repository and demo; Chapter 15 brings its result in as external evidence. When this book says “RELATE” without qualification, it means the corpus.

Embed every item and take the mean cosine similarity within each typed pair:

relation             mean cosine (bge-large)   what it should be
equivalent                 0.96                 high (same claim)
relation-swap              0.99                 LOW  ("Acme acquired Beta" vs "Beta acquired Acme")
partial-support            0.90                 mid
paraphrase                 0.89                 high
entailment                 0.85                 mid-high
negation                   0.83                 LOW  (opposite claim)
contradiction              0.83                 LOW
topic-related              0.77                 mid
temporal-mismatch          0.74                 LOW  (right relation, wrong year)
entity-related             0.72                 low-mid
unrelated                  0.36                 low

Read down the relations that should be low. The one that should stop you is relation-swap: for bge-large, its mean cosine is 0.9883, higher than equivalent at 0.9648. Across all five encoders, relation-swap ranges from 0.9791 to 0.9930 and is higher than equivalent in every model. In RELATE these are constructed asymmetric role reversals: the entities and relation words stay the same, but who did what to whom changes. The vectors barely move. negation (0.83) sits almost on top of contradiction (0.83) and only about 0.06 below paraphrase for bge-large. Across the five models the paraphrase-minus-negation gap runs from about +0.07 down to −0.06 for all-MiniLM-L6-v2 — where the negation is more similar to the source than the paraphrase is. Only unrelated — no shared topic, no shared entity — separates cleanly.

Pause on that 0.99. Cosine is not reporting 99% confidence, and it is not saying the two sentences are 99% equivalent. It is telling you that, under this geometric readout, reversing the argument roles moves the representation almost not at all. If your downstream task cares who did what to whom, high cosine alone is not enough.

MEASURED: the geometry encodes aboutness strongly (paraphrase, topic-related, entity-related all land well above unrelated) and assertion — polarity, argument order, time — barely at all. “X did Y” and “X did not do Y” occupy nearly the same point.

That is not evidence of a broken encoder. It is evidence that generic cosine proximity in these spaces is a poor readout for the assertion-level distinctions RELATE is probing. A plausible mechanism-level reading is that the training signals behind these encoders reward broad semantic relatedness and retrieval usefulness more directly than exact polarity, argument role, or time. Negated and role-reversed sentences therefore retain much of the lexical and topical signal that drives proximity. That interpretation is consistent with what we measured; it is not proof that every embedding objective must behave this way.

The boundary of this result

What the demonstration establishes. For the five encoders tested on RELATE v0.1, generic cosine proximity preserves broad aboutness much more strongly than the assertion-level distinctions probed here. Negation can sit nearly as close as paraphrase, and reversing argument roles produces mean cosines above 0.97 in every tested model.

What it does not establish. Not that embeddings are unreliable, that similarity is useless, or that these exact numbers generalize to encoders outside the sample. “Does this model encode this relationship in a usable way?” is an empirical question with a per-model, per-relation, per-readout answer. Later chapters widen both the sample and the readouts.

That is why the chapter ends with engineering rather than a verdict. If the interpretation of a distance depends on how the space was produced and how it is read, record those conditions.

Lab 1: what does your encoder think “similar” means?

MEASURED — artifact experiments/embeddings-from-first-principles/wave1/artifacts/relation-cosine-by-type.json (bge-large-en-v1.5). REPRODUCIBLE — run_wave1.py 1.1.

Question. For this model, does “similar” mean same-topic, same-claim, or something in between?

What we ran. We embedded all 1,181 typed RELATE pairs with five encoders and took the mean cosine per relation. The table shows bge-large-en-v1.5:

CategoryObserved mean cosDifference from unrelated (0.36)
paraphrase0.89+0.53
negation0.83+0.47
same-topic / different-claim (topic-related)0.77+0.41
entity-overlap only (entity-related)0.72+0.36
unrelated0.36—

The sharpest row is relation-swap: 0.99 for bge-large, higher than the equivalent mean of 0.96. Across all five encoders, the relation-swap mean stays between 0.9791 and 0.9930. The paraphrase-minus-negation gap is only about +0.06 for bge-large, and on all-MiniLM-L6-v2 it is −0.06: the negation sits closer than the paraphrase.

Interpretation. Establishes: this model encodes aboutness strongly and assertion (polarity, argument order, time) barely at all. Does not establish: that every encoder does this — the gap is model-specific, so record your model’s own table in its space record.

A minimal probe you can run yourself:

# pip install sentence-transformers
from sentence_transformers import SentenceTransformer

model = SentenceTransformer("BAAI/bge-large-en-v1.5")

pairs = [
    ("Acme acquired Beta.", "Beta acquired Acme."),  # relation-swap
    ("The trial reported a benefit.", "The trial reported no benefit."),  # negation
    ("In 2018, Acme employed 1,200 people.", "In 2024, Acme employed 8,500 people."),  # temporal mismatch
    ("The cat sat on the mat.", "A cat is sitting on the mat."),  # paraphrase control
]

for a, b in pairs:
    va, vb = model.encode([a, b], normalize_embeddings=True)
    cosine = float(va @ vb)
    print(f"{cosine:.3f}  |  {a}  <>  {b}")

The point is not to reproduce the RELATE table from four invented examples. It is to make the failure mode tangible: hold the model fixed, change one semantic relation at a time, and watch how much the geometry reacts.

Try it yourself

Change one thing at a time and re-run. Swap the encoder for all-MiniLM-L6-v2, replace the pairs with sentences from your own domain, or add several paraphrases as controls. Keep at least one pair whose relationship you are confident about. Then ask: which semantic change barely moves the cosine, and does any negation or role reversal land closer than a faithful paraphrase?

Companion component: the space record

The Embedding Observatory we build by Chapter 26 starts here, with the smallest possible artifact: a record of which transformation produced a vector.

space_record:
  model:          <name>
  version:        <string or commit>
  dimension:      <int>
  normalization:  <none | l2 | whitened>
  pooling:        <cls | mean | last>
  objective:      <what the model was trained to make close>
  notes:          <what "similar" appears to mean, from Lab 1>

This is the rule of thumb from earlier turned into an artifact. You may not know every detail of a model’s training objective, but you can record the model and version, the dimensionality, pooling, normalization, metric assumptions, and what your own probe says “similar” means in this space.

Every vector in the book travels with one of these. A vector without a space record is a list of numbers whose provenance — and therefore much of its usable interpretation — you have chosen to forget.

Failure modes

  • “The embedding captured the meaning.” It captured a geometry shaped by the training process. Ask which relationship that geometry actually preserves before trusting it.
  • Comparing vectors from different models. Two arrays of the same length from different encoders are not in the same space (Chapter 16). The dot product is defined; the interpretation is not.
  • Forgetting normalization. Cosine and dot product agree only when vectors are unit length. Unexpected scores often come from mixing normalized and unnormalized vectors (Chapter 4).
  • Treating one-hot vectors as embeddings. Equidistant identifiers carry no graded structure; nearest-neighbor search over them is exact-match search.

What this chapter established

  • The vocabulary ladder: identifier → feature → vector → representation → embedding → latent representation.
  • The load-bearing claim: an embedding is a learned transformation’s output under an objective, not a container of meaning.
  • Three consequences: the objective shapes what is preserved, the representation is operationally lossy, and the geometry is conditional on model, metric, normalization, and readout.
  • The RELATE corpus and the space record — the two artifacts the whole book reuses.
  • A measured demonstration where negation remains close and a role reversal is even closer than the mean equivalent pair, exposing the gap between aboutness and assertion.

Next

If an embedding is a transformation into a space, then semantic questions become geometric questions in that space — how close, which direction, how dense. The next chapter, Meaning Becomes Geometry, builds a tiny space by hand and watches meaning turn into coordinates, distance, and angle — and marks exactly where that translation starts to leak.