The Shape of an Embedding Space
Freeze the corpus, change only the model, and measure the geometry: common-direction bias, centered-covariance concentration, distance concentration. Then reshape one space by centering and whitening — and find that a near-zero random-pair cosine is not a quality target.
Part II — Inside the Space
Same corpus, five geometries
Freeze the input. Embed exactly the same 1,173 RELATE items with five models and, before scoring any task, measure the geometry each one imposes on identical text.
The evaluation harness requests L2-normalized embeddings (normalize_embeddings=True), so every vector analyzed in this chapter lies on the unit sphere, and cosine is purely a statement about direction.
MEASURED on RELATE v0.1, Wave 2 row 2.8 — artifact
experiments/embeddings-from-first-principles/wave2/artifacts/shape-comparison.json. All descriptors computed from one run on the frozen corpus.
MiniLM-L6 mpnet-base mxbai-large bge-large bge-small
nominal dimension 384 768 1024 1024 384
mean random-pair cosine 0.058 0.075 0.337 0.403 0.451
top-PC variance share (centered) 0.070 0.076 0.087 0.090 0.099
entropy effective rank 259.1 387.4 425.4 434.0 271.0
participation ratio 67.5 67.0 57.0 56.1 48.6
pairwise-distance CV 0.071 0.071 0.087 0.108 0.099
domain-cluster ARI 0.373 0.395 0.408 0.392 0.367
The input texts never changed. Each model brings its own tokenizer, architecture, training history, and output dimension, so the differences below come from the representation pipelines, not from different text samples. This does not isolate any single causal training variable; it shows that “the shape of an embedding space” is really the shape of model + embedding protocol + corpus.
Common-direction bias (row 1) moves by a factor of about eight. For independent draws of unit vectors, E[x · y] = ‖E[x]‖²; because dot product equals cosine on the unit sphere, the random-pair mean is therefore related to the squared norm of the population mean vector. The finite RELATE estimate excludes self-pairs and uses a sample of pairs, so treat it as an empirical common-direction statistic rather than an exact identity for the finite corpus. It runs from 0.058 for MiniLM-L6 to 0.451 for bge-small. For MiniLM and mpnet, the sampled background is centered much closer to cosine zero; for mxbai and the two BGE models, the average sampled pair is already around 0.34–0.45.
This is a background statistic, not “the origin” of cosine space and not a floor (Chapter 5). Cosine zero already means orthogonality. The Chapter 5 anisotropy artifact — from a different random-pair sample — records standard deviations of roughly 0.09–0.12 around the model-specific means. A candidate score of 0.40 therefore occupies very different positions relative to a background centered near 0.06 versus one centered near 0.45, but the mean alone does not make that score “strong,” “weak,” relevant, or noise. Those interpretations require the full background distribution and, ideally, task-specific relevant/irrelevant score distributions and calibration (Chapter 14).
The centered spectrum is a different measurement. The top-PC variance share is σ₁² / Σ σᵢ² computed on the centered matrix V − mean(V) — the mean vector is subtracted before the SVD. So it cannot tell you how large that removed mean vector was. The measured top-PC share is 0.07–0.10 across the five models: the largest centered principal component carries a minority of the residual variance in every case. That is a statement about concentration, not proof of isotropy.
This is worth pausing on, because it is easy to get wrong. A large mean vector and broadly distributed centered covariance are not in tension — they are properties of different objects. bge-small has a mean random-pair cosine of 0.451 and a top centered-PC variance share of 0.099. Both can be true simultaneously because the second statistic is computed only after the mean has been removed.
The centered spectra also differ in other ways. The participation ratio — computed on squared singular values, (Σ σᵢ²)² / Σ σᵢ⁴ — is about 67 for MiniLM and mpnet and about 49–57 for mxbai and the two BGE models, indicating greater variance concentration under that statistic for the latter three on this corpus. Entropy effective rank uses a different weighting convention and also scales with how many spectral directions are available, so raw values from 384-, 768-, and 1,024-dimensional spaces should be interpreted with their ambient dimensions and full spectra in view rather than read as a simple model ordering.
Distance concentration is a third descriptor. The pairwise-distance CV is the standard deviation of sampled pairwise Euclidean distances divided by their mean. All five sit in a fairly narrow band, 0.07–0.11: sampled pairwise distances occupy a comparatively tight range. Nearest-neighbor rankings then turn on comparatively small differences within that band. That is a reason to measure neighbor margins and perturbation stability (Chapter 6), not a reason to call those small differences noise — they may be exactly the task-relevant structure the model learned. And this experiment includes no low-dimensional control, so it establishes that these distributions are concentrated, not that high dimension caused the concentration.
Domain-cluster ARI is not a pure geometric descriptor. It runs k-means on the embeddings and scores the clusters against RELATE’s domain labels with the adjusted Rand index — one coarse external readout of how a clustering aligns with a reference labeling. It barely moves (0.37–0.41) while the common-direction statistic spans an order of magnitude. That does not prove shape is statistically or causally independent of quality; it shows that this one shape statistic is not itself a quality score.
Establishes. On this fixed corpus, the five models produce substantially different common-direction baselines (0.06 to 0.45) and modestly different centered spectra and distance distributions. Geometry is a property of the representation pipeline, measurable before any task.
Does not establish. Which model retrieves better, which captures a relation more faithfully, or which geometry suits an application. Mean random-pair cosine does not fully characterize the directional distribution — a near-zero mean is compatible with unequal covariance eigenvalues, multiple lobes, or low-dimensional structure — so MiniLM and mpnet are “closer to zero than the others on this corpus,” not “isotropic.”
Is a more isotropic-looking space a better semantic representation — and how would we tell?
The shape descriptors, and the hypotheses they raise
Each descriptor says only what it measures. What it motivates is a test, not a diagnosis.
- Common-direction bias. Mean and standard deviation of random-pair cosine, together with the mean-vector norm. For independent unit-vector draws, expected cosine is
‖E[x]‖², so these quantities are closely related but not numerically identical. Motivates: checking whether a raw-score threshold transfers across corpora and model versions (Chapter 14), and inspecting whether a common direction matters to the task. - Centered-covariance concentration. The singular spectrum of
V − mean(V), summarized by the top-PC variance share, the participation ratio (onσ²), and the entropy effective rank (onσ, per Chapter 7’s convention). Motivates: inspecting what a high-variance direction aligns with, then ablating it and scoring the task. - Distance concentration. The coefficient of variation of sampled pairwise distances. Motivates: measuring neighbor margins and perturbation stability (Chapter 6).
- Cluster geometry. The number, size, and separation of dense regions, and how a clustering aligns with a reference labeling. Motivates: checking whether the grouping tracks the property you would claim (Chapters 2 and 6).
A direction whose projection correlates with, say, sentence length could be nuisance for your task, task-relevant, both at once, or a correlate of some third variable. The descriptor cannot tell you which. Only an ablation followed by a task measurement can.
Three interventions, not one
“Reshaping the geometry” covers several distinct operations. Their transforms:
- Centering. Fit the item-space mean vector, subtract it from every item (and from queries under the same fitted transform), then L2-renormalize each result. The subtraction makes the pre-renormalization centered vectors sum to zero. Renormalizing each row can move the final sample mean slightly away from zero again, which is why the post-transform random-pair cosine is measured rather than assumed. This operation removes a fitted common offset before normalization; it does not, by itself, establish that the removed direction was nuisance.
- Dropping top-
kprincipal components. Center, fit PCA on the item corpus, remove the firstkprincipal directions, project onto the remaining subspace, renormalize. Thosekdirections are the highest-variance ones; the experiment does not establish that they carry nuisance. - Whitening. Center, rotate into the PCA basis, divide the retained PCA coordinates by their empirical standard deviations (the implementation adds
1e-8to the denominator), then L2-renormalize each vector. This equalizes the scale of the retained directions before the final per-vector normalization. Because the pipeline ends with L2 normalization, the resulting covariance is not exactly the identity. - Evaluation. Apply the same item-fitted transform to the query vectors and score the downstream task. This step is what decides whether the transform helped the readout.
The Wave 2 artifact tests six variants: raw; centered; drop_top3; whitened_top95var (retain enough leading PCs to explain about 95% of the original centered variance, then divide each projected coordinate by its empirical standard deviation); drop3_whitened_top95var (drop the first 3 PCs, then retain that same count of subsequent PCs and standardize them — this is not “retain 95% of the post-drop variance”); and whitened_full (standardize every centered PCA coordinate). The implementation adds 1e-8 to each standard-deviation denominator and L2-renormalizes afterward. The stored field whitening_gain_ndcg is specifically nDCG(drop3_whitened_top95var) − nDCG(raw).
One framing to carry in: whitening is a deterministic reweighting of directions that already exist. It cannot create semantic information the vectors did not contain. It can change how much of the information already there a fixed readout like cosine exposes — the Chapter 3 and 4 lesson again.
Predict, then measure
Before the numbers, turn the geometric intuition into a prediction. If a large common-direction baseline is hurting cosine ranking, then removing that offset should help most where the baseline is largest. The transform experiment was actually run on three models — MiniLM-L6, mpnet-base, and bge-large — so make the prediction only for those three: which should gain from centering, from dropping the top PCs, and from whitening? Then compare the prediction with the task measurements rather than retrofitting an explanation afterward.
Demonstration: reshaping the RELATE space
MEASURED on RELATE v0.1, Wave 2 row 2.9 — artifact
experiments/embeddings-from-first-principles/wave2/artifacts/whitening-gain.json. Three models (all-MiniLM-L6-v2,all-mpnet-base-v2,bge-large-en-v1.5). Each transform is fit on the RELATE item embeddings and the same fitted transform is applied to the query vectors; retrieval and the hard-negative margin are then scored on RELATE. Fit and evaluation use the same corpus.
mpnet-base nDCG@10 hard-neg margin mean random-pair cosine
raw 0.9518 0.0932 0.0754
centered 0.9523 0.1127 0.0006
drop top 3 PCs 0.9480 0.1060 −0.0010
whiten top-95%-var PCs 0.9098 0.0929 −0.0000
drop 3 then whiten 0.8999 0.0870 −0.0003 ← stored gain: −0.0519
full whitening 0.5269 0.0579 −0.0001
bge-large nDCG@10 hard-neg margin mean random-pair cosine
raw 0.9370 0.0641 0.4030
centered 0.9377 0.0910 0.0017
drop top 3 PCs 0.9341 0.0829 0.0008
whiten top-95%-var PCs 0.8951 0.0611 −0.0009
drop 3 then whiten 0.8922 0.0597 −0.0011 ← stored gain: −0.0448
full whitening 0.0715 −0.0089 −0.0009
MiniLM-L6 nDCG@10 hard-neg margin mean random-pair cosine
raw 0.9357 0.0846 0.0579
centered 0.9323 0.0968 −0.0000
drop top 3 PCs 0.9275 0.0633 −0.0009
whiten top-95%-var PCs 0.9056 0.0736 −0.0008
drop 3 then whiten 0.9000 0.0654 −0.0010 ← stored gain: −0.0357
full whitening 0.8243 0.1043 −0.0001
1. Centering removes almost all of the common-direction bias and barely touches retrieval. For bge-large, the mean random-pair cosine drops from 0.403 to 0.0017 — two orders of magnitude — while nDCG@10 moves from 0.9370 to 0.9377. For mpnet, 0.0754 to 0.0006, with nDCG@10 from 0.9518 to 0.9523. A shape statistic collapsed and the ranking did not. The hard-negative margin actually widened for both (bge-large 0.064 → 0.091; mpnet 0.093 → 0.113). This does not show the mean direction was “nuisance.” It shows that a large change in one geometric property produced almost no change in retrieval ranking quality. Geometric appearance and task usefulness are separable. (MiniLM, whose raw baseline was already 0.058, takes a small centering loss: 0.9357 → 0.9323.)
2. Several very different transforms drive the random-pair cosine to about zero — while retrieval diverges sharply. On mpnet, centering gives nDCG@10 0.9523, top-95%-variance whitening 0.9098, drop-3-then-whiten 0.8999, and full whitening 0.5269; all have mean random-pair cosine within about 0.001 of zero. Reaching the same value on one shape statistic clearly does not imply reaching the same geometry or the same task behavior.
Near-zero random-pair cosine is not a quality target. It says the sampled common-direction statistic is near zero after the transform. It does not establish full isotropy, better semantics, or “cleaner” geometry in any absolute sense.
3. The configured whitening variants reduced nDCG@10 for all three models (stored gains −0.036, −0.045, −0.052). What that establishes is narrow: these fitted whitening interventions did not improve this task metric under this evaluation. It does not, by itself, say why. Candidate explanations include: the original variance weighting was task-useful; whitening over-weighted low-variance directions; the transform changed query and document geometry in an unhelpful way; helpful and harmful effects interacted; the retained-component count was wrong; retrieval simply does not reward more isotropy. Symmetrically — a positive gain would establish that the intervention helped the task, not that any particular direction was nuisance.
4. Full whitening is severely harmful, and by very different amounts. mpnet falls from 0.9518 to 0.5269, MiniLM from 0.9357 to 0.8243, and bge-large from 0.9370 to 0.0715. No random-ranking baseline is reported in this artifact, so do not translate 0.0715 into “near random.” The pipeline standardizes every centered PCA coordinate before the final row normalization. A plausible mechanism is that directions carrying little variance in the fitted corpus receive far more relative weight than before, but the artifact does not identify those directions as noise or prove that this reweighting is the causal mechanism. The measured fact is simpler: full variance equalization badly damaged nDCG@10 for all three models.
5. The hard-negative margin tells a different story from nDCG, which is exactly why both are recorded. Centering widens the margin for all three models: bge-large 0.0641 → 0.0910, mpnet 0.0932 → 0.1127, MiniLM 0.0846 → 0.0968. The configured drop3_whitened_top95var variant lowers the margin below raw for all three, while full whitening is mixed: it lowers the margin for bge-large and mpnet but raises MiniLM from 0.0846 to 0.1043 even as MiniLM’s nDCG falls from 0.9357 to 0.8243. One transform can therefore improve one scalar readout while damaging another. A deterministic transform cannot create information absent from the input vectors, but it can reweight existing information so that a particular metric exposes it differently.
MEASURED: on
mpnet-base,bge-large, andMiniLM-L6/ RELATE v0.1, centering drove the sampled random-pair cosine to about zero with only small nDCG changes. Every configured whitening variant reduced nDCG@10 relative to raw for all three models; the storeddrop3_whitened_top95vargains are −0.0357, −0.0448, and −0.0519. Full whitening was especially damaging to nDCG, while hard-negative-margin effects were stage- and model-dependent rather than uniformly negative.
Why the two headline results are not contradictory. “Centering made mpnet’s random-pair cosine almost zero without materially changing retrieval” and “more aggressive whitening also made it almost zero while badly damaging retrieval” are both true. They do not conflict because random-pair cosine measures one property of the transformed distribution, while nDCG measures whether the resulting ranking still aligns with relevance. Centering subtracts a fitted mean and then renormalizes, so it does change pairwise geometry; empirically, that change happened to leave mpnet retrieval almost unchanged. Whitening makes a much larger intervention by additionally rescaling PCA directions before renormalization. Similar movement in one summary statistic can therefore hide very different changes elsewhere in the geometry.
What this chapter establishes and what it does not
Establishes: a vocabulary that separates common-direction bias, centered-covariance concentration, and distance concentration into distinct measurements; that the five model-plus-protocol combinations produce measurably different shapes on the same RELATE item corpus; and that, for the three models in the transform experiment, centering drove the sampled common-direction statistic close to zero with only small nDCG changes, while every configured whitening variant reduced nDCG@10. The hard-negative margin did not move in lockstep with nDCG, reinforcing that no single readout summarizes task behavior.
Does not establish: that any shape is “correct”; that whitening never helps (older sentence-representation work found gains in different settings — below); that these results transfer to another corpus or to a transform fit on held-out data (fit and evaluation shared the RELATE corpus here); that a shape statistic is independent of task quality; or that a post-hoc transform’s sign reveals anything about how a model was trained.
A historical note on whitening
Methods such as BERT-flow and BERT-whitening improved sentence representations derived from pretrained language models like BERT in the settings their authors studied, where those representations were strongly anisotropic and the readout was cosine similarity or a similar linear comparison. That is a real result in its context. It does not generalize to a rule that whitening usually helps, that it recovers a fixed number of points, or that a non-positive whitening gain marks a model as “already clean” or contrastively trained. The three encoders measured here are modern sentence-embedding models, and this local whitening pipeline did not improve their retrieval — that is the whole claim.
Lab 8: profile five shapes, then reshape three
MEASURED — artifacts
wave2/artifacts/shape-comparison.json(5 models) andwhitening-gain.json(transform pipeline:MiniLM-L6,mpnet-base,bge-large), 1,173 RELATE items, one corpus. REPRODUCIBLE —run_wave2.py 2.8 2.9.
Question. Do different models build different shapes on identical text — and does reshaping a space help the task?
Step 1 — freeze the inputs. The same 1,173 RELATE items for every model.
Step 2 — profile the geometry. Per model: mean and (from the Chapter 5 artifact) standard deviation of random-pair cosine; top-PC variance share on the centered matrix; entropy effective rank; participation ratio; pairwise-distance CV.
| Model | mean random-pair cos | top-PC share | eff. rank | participation ratio | distance CV | domain-cluster ARI |
|---|---|---|---|---|---|---|
| MiniLM-L6 (384-d) | 0.058 | 0.070 | 259.1 | 67.5 | 0.071 | 0.373 |
| mpnet-base (768-d) | 0.075 | 0.076 | 387.4 | 67.0 | 0.071 | 0.395 |
| mxbai-large (1024-d) | 0.337 | 0.087 | 425.4 | 57.0 | 0.087 | 0.408 |
| bge-large (1024-d) | 0.403 | 0.090 | 434.0 | 56.1 | 0.108 | 0.392 |
| bge-small (384-d) | 0.451 | 0.099 | 271.0 | 48.6 | 0.099 | 0.367 |
Step 3 — write a prediction. Which model should gain most from centering? From dropping PCs? From whitening? Commit to it before Step 4.
Step 4 — intervene stage by stage on MiniLM-L6, mpnet-base, and bge-large: raw → centered → drop_top3 → whiten (keep ~95%-var) → drop 3 then whiten → full whitening. Do not skip from raw to whitening.
Step 5 — record geometry and task at every stage: mean random-pair cosine, nDCG@10, and hard-negative margin. Never assess “cleanliness” without the task metric beside it.
mpnet-base bge-large
nDCG@10 margin rand-cos nDCG@10 margin rand-cos
raw 0.9518 0.093 0.075 0.9370 0.064 0.403
centered 0.9523 0.113 0.001 0.9377 0.091 0.002
drop 3 then whiten 0.8999 0.087 −0.000 0.8922 0.060 −0.001
full whitening 0.5269 0.058 −0.000 0.0715 −0.009 −0.001
Step 6 — interpret the dissociation. On mpnet, the random-pair cosine reaches ~0 after centering, after the drop-3-plus-whitening variant, and after full whitening — while nDCG@10 is about 0.95, 0.90, and 0.53 respectively. One shape statistic reaches nearly the same value under three transformations with very different task outcomes. The lesson is not that “more isotropic is worse”; it is that making this one statistic look more isotropic is not sufficient evidence of better retrieval.
Step 7 (PROPOSED — no artifact backs this). Fit the mean, PCA basis, and standard deviations on one RELATE domain split and evaluate on another. A transform that helps the corpus it was fit on may not transfer.
Try it yourself
Add a sixth model. Plot the centered spectra for all six models on one axis, with ambient dimension and normalization/protocol recorded. Sort items by the new model’s top PC once — does it track sentence length, punctuation, or nothing nameable? Then run Steps 4–6 on that model and write one paragraph: its shape, one failure hypothesis the shape raises, and what the measured task deltas per stage actually showed — not what you can infer about its training, which this experiment cannot tell you.
Companion component: the shape profile
The profile records a chain — descriptor, then intervention, then measured consequence, then policy — never descriptor straight to automatic repair.
shape_profile:
space_id:
corpus_hash:
embedding_protocol: # normalization, query/doc prefixes
common_direction:
mean_random_cosine:
std_random_cosine:
mean_vector_norm:
centered_spectrum:
top_pc_variance_share:
effective_rank: # entropy over singular values (Ch 7 convention)
participation_ratio: # over squared singular values
distance_distribution:
sample_size:
mean:
std_over_mean:
external_readouts:
domain_cluster_ari: # a coarse alignment check, not task quality
transform_experiment:
transform_id:
fit_corpus:
eval_corpus:
stages: [raw, centered, drop_top3, whiten_top95, drop3_whiten, full_whiten]
geometry_by_stage: # mean_random_cosine, ...
task_scores_by_stage: # per task: metric, raw score, transformed score
decision:
approved_for_tasks: # tasks where a transformed variant measurably won
rejected_for_tasks: # tasks where it did not
derived_space_id: # a transformed space is a NEW space (Ch 7)
A single nDCG delta does not authorize swapping the production representation. A whitened space is a new derived embedding space with its own provenance and its own calibration, exactly as a compressed space was in Chapter 7. If a transform helps retrieval and hurts classification, one metric cannot approve it globally — the decision belongs to deployment policy weighing every task the space must serve.
Failure modes
- Comparing models by dimension. Tempting because
dis printed on the model card. Check: profile the geometry; two 768-d models can be geometrically unalike, and a 1024-d model can have a lower participation ratio than a 384-d one. - Treating isotropy as the objective. Tempting because a near-zero random-pair cosine looks like a clean space. Check: score the task. On RELATE, three transforms all reached ~0 random-pair cosine with retrieval ranging from unchanged to collapsed.
- Reading a near-zero random-pair cosine as “cleaned” or “good.” Tempting because the number went where you wanted. Check (Chapter 5): it establishes the mean direction was removed — not full isotropy, not better semantics.
- Calling a length-correlated PC a “nuisance axis inflating scores.” Tempting because the correlation is easy to compute. Check: inspect the direction, form a hypothesis, ablate it, then score the task. Only the last step tells you whether removing it helps.
- Reading a whitening gain’s sign as a diagnosis of the raw geometry. Tempting because “it helped / it hurt” feels explanatory. Check: the sign says whether this intervention helped this metric under this evaluation. It does not identify which direction was nuisance, or whether the raw geometry was “clean,” or how the model was trained.
- Letting one task metric authorize a derived space. Tempting because you measured one number and it moved. Check: a transformed space needs a decision across every task it will serve, and its own calibration (Chapter 14).
What this chapter established
- The empirical object is
shape_profile(model, embedding protocol, corpus)— not a context-free property of a model. Common-direction bias, centered-covariance concentration, and distance concentration describe different aspects of the distribution, and a large common mean direction is fully compatible with broadly distributed centered variance. - A near-zero random-pair cosine is not a quality target. On mpnet, centering, drop-3-plus-whitening, and full whitening all drove it to ~0 while nDCG@10 landed at roughly 0.95, 0.90, and 0.53.
- Every configured whitening variant reduced nDCG@10 for all three tested models, while the hard-negative margin moved independently of nDCG — one transform can improve one readout and damage another.
- A deterministic post-hoc transform cannot invent information its inputs lack, but it can discard information or re-weight what a fixed readout exposes. Its task delta tells you what happened under that evaluation, not why — a shape statistic identifies a property worth investigating; only task evidence says whether changing it helps.
Next
Part II characterized the space at rest. Part III puts it to work, and treats retrieval as an experiment. The next chapter builds retrieval from its primitive operation — embed, compare, rank — in a handful of lines, with no vector database and no approximate search, so every parameter of that operation is visible before any index is introduced.