← Cellular Automata From First Principles

Test Generalization Beyond Training

A model can look robust while still depending on the exact conditions used during training.

So after growth, persistence and regeneration, we need a harder question:

what happens when the world changes?

Build a generalization matrix

Vary dimensions independently:

canvas size
seed position
update rate
rollout length
damage geometry
damage severity
noise level
boundary conditions

Then evaluate combinations that were not used during training. The table at the end of this chapter is executable scaffolding, not a report: every row specifies how to measure, and the verdicts get filled in by running, starting from one held-out axis at a time rather than all shifts at once.

def evaluate_condition(model, make_initial, target, steps=128):
    state = make_initial()
    final = rollout(model, state, steps)
    return float(F.mse_loss(final[:, :4], target))

Shift the seed

If training always starts at the center, test elsewhere:

def make_seed_at(y, x, size=96, channels=16):
    state = torch.zeros(1, channels, size, size, device=DEVICE)
    state[:, 3:, y, x] = 1.0
    return state

(Same convention as Chapter 39’s seed: alpha and hidden set, RGB dark — one seed semantics across Part V.)

Because the rule is local and shared spatially, translation should be a natural capability when boundaries do not interfere. But we should measure it, not assume it.

Change the canvas size

Train on 64×64 and evaluate on larger grids:

seed = make_seed_at(48, 48, size=96)

If the target still develops correctly, that is evidence the system has not simply encoded one fixed array position.

Inject state noise

def add_state_noise(x, sigma=0.02):
    return x + sigma * torch.randn_like(x)

Test several noise levels and report performance curves rather than one anecdotal example.

Change update rates

A model trained around a 50% firing rate may fail at 20% or 90%.

rates = [0.2, 0.35, 0.5, 0.65, 0.8, 1.0]

This tests whether local coordination survives a different effective timescale.

Hold out perturbations

If training uses circular wounds, evaluate rectangles and slices.

If training uses small wounds, evaluate larger ones.

This separates:

robustness to familiar corruption

from:

robustness to new corruption

Report a table, not a victory image

For example — template only; every entry below must be measured, including the failures:

condition             final loss   recovery time   survived
-------------------------------------------------------------
center seed           ...          ...             ...
shifted seed          ...          ...             ...
96x96 canvas          ...          ...             ...
20% fire rate         ...          ...             ...
large slice damage    ...          ...             ...
noise sigma=0.05      ...          ...             ...

This is much more informative than selecting the best animation — provided “survived” means measured against the task threshold, not eyeballed.

Generalization has a boundary

A local learned rule may generalize impressively within one family of dynamics while failing abruptly outside it.

That boundary is scientifically interesting.

Do not hide it.

Map it. And keep the chapter’s central doctrine attached to every map:

success outside the training example is still not success outside the training distribution.

works on held-out examples
  ≠ broad generalization

works on nearby conditions
  ≠ out-of-distribution robustness

regeneration on unseen damage
  ≠ arbitrary repair

The next transition

We now have training setups and test protocols for four capabilities, each of which has to be demonstrated rather than assumed:

grow
persist
repair
survive some distribution shift

So far, generalization has meant preserving a learned form under changed conditions. There is a harder test. Instead of asking the automaton to preserve a form, we can ask local updates to carry out a computation whose answer depends on information distributed across the grid.

That is distributed computation over a grid, and the same architecture can be pointed at it.

In the next chapters we will use NCA for pathfinding and maze-like computation, then inspect what the learned system is actually doing internally.


Research

  • Earle, S., Yildiz, O., Togelius, J. & Hegde, C. — Pathfinding Neural Cellular Automata (2023). The maze-domain precedent for holding out whole axes: hand-coded and learned BFS/DFS NCAs, diameter computation with strong generalization, and adversarially evolved training mazes that improve out-of-distribution robustness. The standard this chapter’s matrix is measured against. https://arxiv.org/abs/2301.06820

  • Mordvintsev, A., Randazzo, E., Niklasson, E. & Levin, M. — Growing Neural Cellular Automata (Distill, 2020). The distribution-expansion methods this chapter generalizes: sample pools and damage-augmented training as the in-distribution widening that precedes genuine held-out testing. https://doi.org/10.23915/distill.00023