← Embeddings From First Principles

How Many Dimensions Does Meaning Need?

One 768-dimensional representation, described by numbers from about 4 to 768 depending on what you measure. Separate ambient size, spectral size, local intrinsic-dimension estimates, and task-retention thresholds — and discover that no single whole-space geometry number supplies the task-safe compression dimension.

Part II — Inside the Space

A representation with several sizes

all-mpnet-base-v2 emits vectors of length 768. Embed the 1,173 RELATE items with it and ask several different questions about the size of the resulting cloud, and you do not get one answer:

nominal dimension                        768
centered matrix-rank ceiling             768   (= min(1,173−1, 768))
95%-of-variance dimension                205
participation ratio                       67.0
entropy effective rank                   387.4
intrinsic-dimension estimate (MLE, k=10)   6.81
intrinsic-dimension estimate (TwoNN)       4.23

The Wave 2 dimensionality artifact records the last five numerical summaries plus nominal dimension; it does not persist an observed matrix-rank value, so the line above reports only the algebraic ceiling. Same vectors, same corpus: the reported or implied notions of size still range from about 4 to 768 — almost two orders of magnitude.

The tempting question is “which one is the real dimension?” That is the wrong question. Each of these numbers answers a different one, and reporting any single value as “the dimensionality” throws away the fact that they measure different things. And the number an engineer actually needs — how many dimensions this representation can lose before a specific task degrades — is a fourth quantity again, not on that list.

Which notion of dimensionality are we asking about, and does any of them predict how far a given task can be compressed?

Four layers of “how big”

It helps to sort the numbers into a small hierarchy. The first three layers describe geometry. The fourth answers an engineering question, and it is the one the chapter’s experiment is built around.

1. Ambient and algebraic size

Nominal dimension d is the length of the array — a model specification, not a measurement of the data.

Matrix rank is the number of linearly independent directions present in the centered sample matrix, equivalently the number of non-zero singular values in exact arithmetic. Centering makes the row vectors sum to zero, so rank(X_centered) ≤ min(n − 1, d). For RELATE/mpnet, that ceiling is min(1,172, 768) = 768, so centering does not force the rank below the model width. By contrast, a 1,536-wide representation measured on only 1,173 items could have centered sample rank at most 1,172. Numerical rank also depends on a tolerance: tiny singular values can be non-zero without being practically important. Rank is therefore partly a statement about the sampled matrix, not a direct measure of semantic capacity.

2. Spectral size

This layer asks how the centered variance is spread across orthogonal directions — the singular-value spectrum σ₁ ≥ σ₂ ≥ ….

  • 95%-variance dimension. The smallest number of leading principal components whose squared singular values account for at least 95% of the total centered variance. The 95% is a chosen threshold, not a natural constant; at 99% the RELATE/mpnet figure rises from 205 to 352.
  • Entropy effective rank. With pᵢ = σᵢ / Σⱼ σⱼ — normalized singular values, not variances — the effective rank is exp(−Σ pᵢ log pᵢ). If the spectrum were flat across r directions this returns r; the more the spectrum concentrates, the lower it drops. Conventions differ: some authors take the entropy over eigenvalues (variances, σᵢ²) instead. This book’s artifact uses the singular-value convention, and the measured value 387.4 depends on that choice.
  • Participation ratio. With λᵢ = σᵢ², the participation ratio is (Σ λᵢ)² / Σ λᵢ². Like effective rank it returns r for a flat spectrum and falls as variance concentrates, but it is computed on the squared singular values. Squaring exaggerates the lead of the top directions, so the participation ratio (67.0) comes out well below the entropy effective rank (387.4) on the same spectrum. Neither is more “correct”; they weight the spectrum differently.
  • Stable rank, ‖X‖_F² / ‖X‖₂² = Σ σᵢ² / σ₁², is another inverse-concentration measure. It was not computed in this book’s dimensionality artifact, so it appears here only as a concept, not a measured value.

3. Local geometric size

Intrinsic-dimension estimators try to infer the number of local degrees of freedom from how neighbor distances behave.

  • The TwoNN estimate (Facco et al., 2017) uses the ratio of each point’s second- to first-nearest-neighbor distance. On RELATE/mpnet: 4.23.
  • The MLE estimate (Levina & Bickel, 2004) uses the distances to the first k = 10 neighbors. On RELATE/mpnet: 6.81.

Intrinsic dimension is not directly observed. These estimators are sensitive to sampling density, duplicates, the distance metric, curvature, noise, scale, and their own modeling assumptions — which is why the two disagree here. Write “the TwoNN estimate is 4.23,” never “the intrinsic dimension is 4.2.” And, keeping Chapter 6’s discipline: a low local-ID estimate is compatible with lower-dimensional local structure, but it does not establish that RELATE lies on a smooth manifold of dimension 4 to 7. RELATE’s items are built from templates, which plausibly flattens local patches; that is a hypothesis about why the estimates are low, not a measured fact about a surface.

4. Task-retention size

The smallest compressed dimension, within a tested grid, that keeps a specified task metric within a chosen tolerance of a reference condition. This is not a property of the point cloud. It is the answer to “am I allowed to compress, and by how much” — a policy question — and it has to be measured against the task you actually run. Chapter 6’s rule applies unchanged: a geometric detector does not authorize an intervention.

What could produce a spread this wide

The gap between ~4 and ~768 is real. Its causes are not cleanly separable from these measurements alone. Plausible contributors:

  • Correlated variation. If variation is correlated, the covariance spectrum concentrates and PCA can summarize much of that variation with fewer data-defined directions.
  • Directional nonuniformity (Chapter 5). A nonuniform vector distribution changes the spectrum and can concentrate variance into preferred directions.
  • Lower-dimensional local structure. Neighborhoods may have fewer effective degrees of freedom than the ambient dimension.
  • Architecture, objective, and pooling. Training and representation choices determine which distinctions are exposed and how their variance is distributed; there is no requirement that information be spread uniformly across all d coordinates.
  • Corpus composition. The measured spectrum and local-neighbor statistics are properties of the model on this sampled corpus. A broader or differently composed corpus can move either estimate up or down; the direction has to be measured.
  • Finite sampling. With 1,173 points, high-dimensional spectral and local estimates should be read as sample-dependent summaries rather than population constants.

This chapter measures the spread. It does not decompose it into causes.

What the numbers do and do not license

  • Redundancy relative to a measurement is common in modern text-embedding settings — but “every model wastes most of its dimensions” is stronger than the evidence supports. The defensible claim is that, for many tasks and models, a representation carries more nominal dimensions than that task needs at a stated tolerance.
  • Low-variance directions are not “unused.” Chapter 5 showed that coordinates are basis-dependent, and low-variance principal directions can still carry task-relevant signal. Prefer “variance is concentrated” or “redundant relative to task X at tolerance t” over “the other dimensions do nothing.”
  • Effective rank is a model-comparison axis, not a verdict. On RELATE, minilm-l6 has effective rank 259 out of 384, mpnet-base 387 out of 768, bge-large 434 out of 1024. Which model is “using its capacity” is a question about task performance (Chapters 8 and 13), not one the spectrum answers on its own.

Compression operators are not one thing

People say “truncation” for several different operations. Each produces a new derived space with its own identity (Chapter 17), and they degrade differently.

operationwhat it doesnotes
Raw-prefix croppingkeep the first k native coordinates of an ordinary modelbasis-dependent (Chapter 5); only meaningful if the model was trained to make prefixes useful
Post-hoc PCA truncationfit PCA on a corpus, rotate into principal directions, keep the top kkeeps data-defined directions, not arbitrary coordinates; the PCA basis rotates with the data, so this is not the same operation as cropping raw coordinates
Matryoshka-prefix truncation (Kusupati et al., 2022)keep the first k coordinates of a model trained with a nested multi-scale loss“useful by construction” describes the training objective’s intent; actual retention at a given k still has to be measured
Random projectionapply a fixed random linear map to k dimensionssee Johnson–Lindenstrauss, below
Learned projectiontrain a task- or data-specific map to k < d dimensions (for example an encoder/projection head or specialized compression method such as SMEC)optimized for a chosen objective rather than fixed by variance or randomness; dimensional reduction is not one-to-one over the full input space, and task retention still has to be measured

Johnson–Lindenstrauss, precisely

The JL lemma states that for a fixed set of n points, a suitable random linear projection into a dimension on the order of ε⁻² log n can preserve every pairwise Euclidean distance within a multiplicative factor 1 ± ε, with high probability (the exact constants and probability terms depend on the theorem variant). That is a distance-distortion guarantee, not a PCA statement and not a task-quality guarantee. It also does not automatically preserve nearest-neighbor order: if two original distances d₁ < d₂ are close enough that their allowed distortion intervals overlap — for example when (1 + ε)d₁ ≥ (1 − ε)d₂ — the farther candidate can legally become the nearer one after projection. Retrieval can depend on exactly those small relative margins. Treat JL as a sufficient asymptotic scaling result for random projections, not a measured optimum; the random-projection runs later in this chapter score downstream tasks rather than testing JL’s all-pairs distortion guarantee.

Compressibility is a property of the task

The clearest external evidence is Tsukagoshi & Sasano (2025), on prompt-based text embeddings. Their experiments find substantial redundancy overall, and — the directionally important part — classification and clustering tolerate far more aggressive dimensionality reduction than retrieval and semantic textual similarity. Because task-specific prompts change the representation itself, this is not a single fixed spectrum read four ways; the intrinsic dimensionality and isotropy differ across the task-specific representations. The lesson that transfers: how much you can compress is a property of the task and its readout, not of the model alone.

Geometry can suggest redundancy. Only a task-preservation experiment can authorize compression.

Demonstration: RELATE’s retention curves

MEASURED on RELATE v0.1, Wave 2 rows 2.2–2.7 — artifacts under experiments/embeddings-from-first-principles/wave2/artifacts/. Primary model all-mpnet-base-v2 (768-d). STS uses a RELATE-native relation→similarity proxy (dim_common.STS_TARGET), not human labels.

The several sizes, measured (row 2.2).

nominal dimension                        768
95%-of-variance dimension                205
participation ratio                       67.0
entropy effective rank                   387.4
intrinsic-dimension estimate (MLE k=10)    6.81
intrinsic-dimension estimate (TwoNN)       4.23

Row 2.2 does not persist matrix rank. The centered sample-rank ceiling is 768 here (min(1,173−1, 768)), but that algebraic bound is not an additional measured artifact value.

The hypothesis. Perhaps one of those geometry numbers predicts how far this representation can be compressed before a task breaks.

The retention experiment, defined precisely (row 2.3).

  • The PCA basis is fit on all 1,173 RELATE item vectors. Those same item vectors are used by the item-based evaluations, while retrieval and hard-negative scoring also transform RELATE query vectors through the fitted PCA map. There is no held-out corpus for fitting the projection here; throughout this section, “retained within tolerance” means under this measured RELATE evaluation, not “validated to generalize to a new corpus or domain.”
  • The pipeline centers the vectors, rotates them into the PCA basis, keeps the top d components, and renormalizes each to unit length — at every d, including d = 768. So the d = 768 endpoint is a processed reference, not the untouched model output. That is why full-dimensional PCA retrieval scores nDCG@10 0.9523 here, slightly above the raw-model 0.9518 from Chapter 4.
  • Four task scorers, each higher-is-better: retrieval nDCG@10 (graded), an STS proxy (Spearman of cosine against the relation→similarity map), clustering ARI (k-means against domain labels), and the hard-negative margin (mean of the grade-3 positive’s score minus the best structured-perturbation hard negative’s score).
  • The 5%-retention threshold is the smallest tested d whose task score is at least 0.95 × the processed d = 768 reference. This is an operational threshold tied to the tested grid and the reference — not an automatically detected elbow.
d       retrieval   STS proxy   clustering ARI   hard-neg margin
768      0.9523      0.6344       0.3922            0.1127    ← processed reference
256      0.9528      0.6393       0.3922            0.1132
128      0.9571      0.6435       0.3924            0.1166
 64      0.9476      0.6335       0.4319            0.1216
 32      0.9266      0.6072       0.4376            0.1316
 24      0.9129      0.5760       0.4347            0.1414
 16      0.8740      0.5495       0.4389            0.1270
  8      0.7218      0.5247       0.4351            0.1287
5%-retention threshold:  24         32           8               8

The curves are not monotone. Retrieval rises slightly above the processed full reference around d = 96–192 (0.9563–0.9571 versus 0.9523) before falling. Clustering ARI rises from about 0.392 to about 0.44 under aggressive reduction, and the hard-negative margin rises from about 0.113 to a peak of 0.141 at d = 24. Those improvements are measured outcomes, not a causal diagnosis: PCA may remove variation irrelevant or harmful to a particular scorer, but this experiment does not identify which directions caused the change or why. The broad pattern is “stable or sometimes better over a range, then worse under sufficiently aggressive reduction,” and the bumps are exactly why the retention threshold must be defined operationally rather than read off a visually smooth elbow.

Four tasks, four thresholds: clustering 8, hard-negative discrimination 8, retrieval 24, the STS proxy 32 — a spread of 24 on a single model. Each threshold is the smallest tested dimension that qualifies under the 5% rule. For retrieval, the next tested dimensions below 24 fail the criterion (d = 16 gives 0.8740; d = 8 gives 0.7218). For clustering and hard-negative margin, 8 is the floor of the tested grid, so the experiment says nothing about whether 4, 6, or 7 dimensions would still qualify. A threshold at the grid floor is therefore a bound from the experiment, not proof that 8 is minimal in the continuum.

No geometry number is the threshold (row 2.4).

threshold ÷ TwoNN-ID estimate (4.23)       1.9×  (clustering, hard-neg)  …  7.6×  (STS)
threshold ÷ MLE-ID estimate  (6.81)        1.2×                          …  4.7×
threshold ÷ entropy effective rank (387.4) 0.02×                         …  0.08×
threshold ÷ 95%-variance dimension (205)   0.04×                         …  0.16×

None is a constant, and none lands on a threshold. The smallest dimension the grid tested, d = 8 — about 1.9× the TwoNN estimate and 1.2× the MLE estimate — already costs retrieval 23 nDCG points (0.9523 → 0.7218). The experiment never ran d = 5 or d = ID; d = 8 is the floor of the grid. The lesson is not “PCA cannot reconstruct a curved manifold” — that is not established here. It is narrower: a local intrinsic-dimension estimate answers a different question from “how many global linear PCA components preserve this retrieval task.”

A candidate predictor, still a hypothesis (row 2.5). The entropy effective rank of the between-relation scatter — the covariance of mean pair-difference vectors grouped by the 11 RELATE relation labels — is 5.42, numerically close to the two measured 8-dimensional thresholds and far below the whole-space effective rank of 387.4. But the analogous retrieval-restricted rank (from the covariance of query-minus-relevant-centroid directions) is 74.76, about 3.1× the retrieval threshold of 24. Two cautions dominate the interpretation: the between-relation construction uses only about 11 group means, which structurally limits the rank of the object being summarized, and the candidate was constructed and assessed on the same experimental system with no held-out replication or chance/baseline comparison. Treat 5.42 as an exploratory BOOK HYPOTHESIS worth trying to falsify, not as evidence that a task-restricted rank predicts compression thresholds.

Operator comparison, with the confound marked (rows 2.6–2.7). On the same model, mpnet-base: post-hoc PCA has a retrieval threshold of 24; a single seeded random Gaussian projection has a threshold of 48. At every matched dimension from 16 to 512, PCA’s retrieval nDCG and hard-negative margin exceeded the sampled random projection’s (at d = 16, 0.874 versus 0.756; at d = 64, 0.948 versus 0.910). Only one random projection was drawn per dimension, so this is “PCA beat the sampled random projections in this experiment,” not a statement about the distribution of random-projection outcomes. The Matryoshka-prefix arm used a different model — mxbai-embed-large-v1, 1,024-d, MRL-trained — with a retrieval threshold of 64. Model and operator are confounded there, so treat it as a separate native-model demonstration, not as evidence that PCA is intrinsically better than Matryoshka.

MEASURED: on mpnet-base / RELATE v0.1, task quality is broadly stable over a range of PCA dimensions and then degrades under aggressive reduction, with non-monotone curves. The 5%-retention threshold is task-set — 8, 8, 24, 32 on one model — and does not equal any spectral summary (387.4, 67.0, 205) or intrinsic-dimension estimate (6.81, 4.23). PCA outperformed a single seeded random Gaussian projection at every matched dimension on this model. The operator comparison against Matryoshka is confounded by base model.

What this chapter establishes and what it does not

Establishes: nominal dimension, algebraic/sample-rank bounds, 95%-variance dimension, entropy effective rank, participation ratio, and the TwoNN and MLE intrinsic-dimension estimators are distinct quantities answering distinct questions. On mpnet-base / RELATE, the persisted geometric summaries span roughly 4 to 768. The 5%-retention threshold is another quantity again, measured per task (8 / 8 / 24 / 32); none of the tested whole-space summary values supplies one task-independent compression threshold. In the same-model comparison, post-hoc PCA also outperformed the one sampled Gaussian random projection at every matched reported dimension.

Does not establish: a “true” dimensionality for the space; a universal target storage dimension; that any measured threshold transfers to another corpus, model, or split (the PCA basis was fit and evaluated on the same RELATE items); that the task-restricted effective rank predicts thresholds (one system, no holdout, no chance baseline); that PCA is intrinsically superior to Matryoshka (different base models); that one random projection per dimension characterizes random-projection behavior; that a low intrinsic-dimension estimate implies mode collapse; or that low-variance directions carry no task-relevant signal.

Lab 7: measure the real size of your space, then test what predicts compression

MEASURED — artifacts wave2/artifacts/dimensionality-report.json, retention-curves.json, knee-vs-id.json, task-restricted-effrank.json (mpnet-base, 1,173 RELATE items). REPRODUCIBLE — run_wave2.py 2.2 2.3 2.4 2.5.

Question. How many dimensions does your space use — and does any single number predict how far you can compress it?

Step 1 — measure the geometry. SVD the centered embedding matrix (RELATE v0.1: 1,173 item vectors, nominal d = 768). Report: entropy effective rank 387.4, participation ratio 67.0, 95%-variance dimension 205, MLE-ID estimate 6.81, TwoNN-ID estimate 4.23.

Step 2 — predict before you sweep. Write down your task, your metric, and your tolerance now. From each candidate statistic, predict the compressed dimension you expect to be safe: effective rank, participation ratio, 95%-variance dimension, 2 × TwoNN-ID, the TwoNN-ID itself.

Step 3 — run the per-task PCA retention sweep.

d keptretrieval nDCG@10STS proxy ρclustering ARIhard-neg margin
768 (processed reference)0.95230.63440.39220.1127
1280.95710.64350.39240.1166
640.94760.63350.43190.1216
240.91290.57600.43470.1414
8 (grid floor)0.72180.52470.43510.1287

5%-retention thresholds: clustering 8, hard-negative 8, retrieval 24, STS proxy 32 — spread 24.

Step 4 — see which predictions survived. threshold ÷ TwoNN-ID estimate spans 1.9× to 7.6× across tasks; threshold ÷ effective rank spans 0.02× to 0.08×. No whole-space number is the threshold. The one lead: the between-relation-scatter effective rank (5.42) lands near the relation-separation thresholds (8), while the retrieval-restricted version (74.76) overshoots the retrieval threshold (24) by 3×.

Try it yourself

Extended experiments; no artifact backs these yet.

  • Hold out a domain. Fit the PCA basis on the other RELATE domains and evaluate on the held-out one (_retention_curves(..., domain_holdout=...) supports this). A threshold measured with fit and evaluation on the same corpus is not evidence of generalization.
  • Add a second model. Re-run Steps 1–4 with d ∈ {full, 256, 128, effRank, 2·ID, ID}. Report each threshold and its ratio to (effective rank, TwoNN-ID, 2×ID, 95%-variance dim).
  • Repeat the random projection with several seeds and report the spread, rather than trusting one draw.
  • Test the hypothesis properly. Replication of the task-restricted-effective-rank lead on a second split and a second model is the test of whether it is a predictor or a coincidence.
  • Record the retained space’s transform, seed (if any), fit split, eval split, task, metric, tolerance, and hash — a bare “safe d = 24” carries none of the conditions that made 24 safe.

Companion component: the dimensionality report

The report is a record of evidence, not a dimension calculator. Every field carries the context that makes its number interpretable.

dimensionality_report:
  space_id:
  corpus_id:

  ambient:
    nominal_d:
    matrix_rank:                     # bounded by min(n-1, d)

  spectrum:
    definition:                      # e.g. "entropy over normalized singular values"
    effective_rank:
    participation_ratio:
    variance_dim_95:
    singular_values:

  intrinsic_dimension:               # a list — the estimators disagree
    - estimator:                     # twonn | mle | ...
      parameters:                    # e.g. k=10, metric=euclidean
      estimate:

    retention:                         # one row per task, from a measured curve
    - task:
      metric:
      tolerance:
      tolerance_rule:                # e.g. score(d) >= 0.95 * reference_score
      reference_condition:
      reference_score:
      compression_method:            # raw_prefix | pca_posthoc | matryoshka_prefix | random_projection | learned
      fit_split:
      eval_split:
      tested_dimensions:
      scores_by_dimension:
      smallest_qualifying_d:

  transform:
    derived_space_id:
    transform_hash:                  # fitted PCA basis / projection identity
    input_normalization:
    output_normalization:
    random_seed_if_any:

A bootstrap or resampling interval on the intrinsic-dimension estimate is a recommended extension, not a field to fill with an invented number. And a low effective-rank / nominal-d ratio is a geometry observation, not a defect — it may reflect the corpus, the objective, or useful concentration, and it means nothing about wasted capacity until task evidence is attached.

The governing rule, echoing Chapter 6: dimensionality statistics are diagnostics; compression is a policy. A low intrinsic-dimension estimate is not permission to truncate. A low effective rank is not permission to truncate. A per-task retention curve against the metric you actually run is the evidence that can authorize a specific compression choice — and the resulting space is registered as a new derived space (Chapter 17) with its own hash, metric configuration, and calibration.

Failure modes

  • “It’s a 1,536-dimensional model.” Confuses array width with measured geometry. Check: measure the sampled spectrum and local geometry on the corpus you actually use, then run task-retention curves. The matrix may be numerically full rank while its spectral summaries and task-retention thresholds are much smaller; those are different statements.
  • “The ID estimate is 5, so compress to five dimensions.” Confuses a local geometric estimate with a global linear task-preservation requirement. Check: run the retention curve — on RELATE, d = 8 (already above both ID estimates) cost retrieval 23 nDCG points.
  • “95% of variance kept means 95% of task quality.” Confuses variance reconstruction with task performance. Check: score the task at each dimension, not the reconstructed variance.
  • “PCA beats Matryoshka.” Confuses an operator comparison with a model-plus-operator comparison. Check: compare operators on the same base model; the book’s Matryoshka arm used a different one.
  • “Random projection lost, so JL failed.” Confuses a task-metric experiment with a theorem about pairwise-distance distortion on a fixed point set. Check: JL guarantees distance preservation, not ranking preservation or task quality.
  • “Compression improved clustering, so the dropped dimensions were noise.” Turns an outcome into an unsupported causal claim. Check: the ARI-versus-domain score rose; the mechanism — regularization, removal of label-irrelevant variance, something else — was not identified.
  • “Safe dimension = 24.” Strips the task, metric, tolerance, tested grid, transform, fit corpus, and evaluation split that made 24 the answer. Check: record all of them, or the number does not travel.

What this chapter established

  • “Dimensionality” is not one quantity. Nominal width, sample-matrix rank, 95%-variance dimension, entropy effective rank, participation ratio, and the TwoNN/MLE estimators are different mathematical objects answering different questions — and intrinsic dimension is estimated, not observed, which is why the two estimators here disagree with each other.
  • The task-retention threshold is a quantity distinct from all of them: an operational “smallest tested d within tolerance,” tied to a task, a metric, a tested grid, a projection pipeline, and a reference. No whole-space summary supplied it, and the retention curves were not even monotone.
  • Raw-prefix cropping, post-hoc PCA, Matryoshka-prefix truncation, random projection, and learned reduction are distinct transformations producing distinct derived spaces. Johnson–Lindenstrauss is a distance-distortion guarantee for a fixed point set — not a PCA spectrum, not a task-safe dimension, and not automatic nearest-neighbor-order preservation.
  • Governing rule: geometry can suggest redundancy; only a task-preservation experiment can authorize compression.

Next

We now know that nominal width, the spectral summaries, and the task-retention threshold are different numbers. The next chapter asks what shape the remaining variance has: how anisotropic the cloud is, how concentrated the spectrum, which directions dominate, how tightly distances cluster — and how sharply those properties differ across real models measured on the identical corpus.