Make the Automaton Differentiable
So far every rule in this book has been chosen by us.
Even Lenia still asks us to decide the neighborhood kernel, the growth function and the parameters that connect them.
What if we stop designing the local rule directly?
What if we define only the goal, then let gradient descent discover a local update rule that achieves it? (Designed here, learned starting now — Chapter 36’s bridge sentence becomes this chapter’s premise.)
That requires one major change:
the cellular automaton must become differentiable.
A cellular automaton is already a repeated function
Every system we have built can be written as:
state_t
↓
local perception
↓
local update rule
↓
state_t+1
Or mathematically:
x_(t+1) = F(x_t)
If F contains trainable parameters theta:
x_(t+1) = F_theta(x_t)
then after many steps:
x_T = F_theta(F_theta(...F_theta(x_0)))
If those operations are differentiable, a loss measured at x_T can send gradients all the way back into theta.
The automaton becomes a recurrent neural system.
Start with PyTorch tensors
import torch
import torch.nn.functional as F
DEVICE = "cuda" if torch.cuda.is_available() else "cpu"
state = torch.zeros(1, 1, 64, 64, device=DEVICE)
state[:, :, 32, 32] = 1.0
The dimensions are:
batch
channel
height
width
This is already a useful change from our earlier NumPy examples because PyTorch can record the operations used to transform the state.
Note the channel count: this chapter’s state carries one visible channel only. The full neural cellular automaton carries sixteen — RGB, an alpha “alive” channel, and twelve hidden channels with no predefined meaning. Hidden channels arrive in Chapter 39; until then, every value here is directly rendered.
Differentiable neighborhood perception
A classical automaton might explicitly count neighbors.
A differentiable automaton can perceive its neighborhood with convolution.
kernel = torch.tensor(
[
[0.0, 1.0, 0.0],
[1.0, 0.0, 1.0],
[0.0, 1.0, 0.0],
],
device=DEVICE,
).view(1, 1, 3, 3)
perception = F.conv2d(state, kernel, padding=1)
Nothing here requires a hard decision such as:
if neighbors == 3:
become alive
The perception is a real-valued tensor that can flow into smooth functions.
(The canonical NCA perceives with fixed Sobel gradient filters — a design choice, not a requirement. This chapter’s cross kernel teaches the mechanism with less machinery; the Sobel variant belongs to the full architecture in Chapter 38.)
Replace hard thresholds with smooth functions
A hard threshold:
alive = (perception > 2.5).float()
is awkward for gradient-based learning because the output changes abruptly.
Instead we can use a sigmoid:
alive = torch.sigmoid(8.0 * (perception - 2.5))
The output still behaves like a threshold, but now small changes in the input create small changes in the output.
That creates useful derivatives.
A trainable local rule
Let us make the threshold itself learnable.
threshold = torch.nn.Parameter(torch.tensor(2.5, device=DEVICE))
sharpness = torch.nn.Parameter(torch.tensor(5.0, device=DEVICE))
def step(state):
perception = F.conv2d(state, kernel, padding=1)
return torch.sigmoid(sharpness * (perception - threshold))
(step keeps its book-wide name deliberately: discrete transition, continuous integration, and now differentiable update are three mechanisms for one role — the automaton’s transition. The signature narrows each time; the loop around it does not change.)
Now the local rule has parameters.
We can optimize them.
Define a target
Suppose we want the automaton to produce a ring.
y, x = torch.meshgrid(
torch.arange(64, device=DEVICE),
torch.arange(64, device=DEVICE),
indexing="ij",
)
r = torch.sqrt((x - 32) ** 2 + (y - 32) ** 2)
target = ((r > 10) & (r < 14)).float()[None, None]
The loss can simply compare the final state to the target:
def loss_fn(state):
return F.mse_loss(state, target)
Unroll the automaton
def rollout(initial_state, steps):
state = initial_state
for _ in range(steps):
state = step(state)
return state
rollout is now a load-bearing vocabulary word: executing a fixed-parameter rule for N steps, with no learning inside. Training updates parameters; rollout merely runs them. Everything called “deployment” later is rollout with frozen weights.
Then train:
optimizer = torch.optim.Adam([threshold, sharpness], lr=1e-2)
for iteration in range(1000):
initial = torch.zeros_like(target)
initial[:, :, 32, 32] = 1.0
final = rollout(initial, steps=20)
loss = loss_fn(final)
optimizer.zero_grad()
loss.backward()
optimizer.step()
(Verified on CPU at small scale: shapes consistent, gradients reach both parameters, parameters move, loss decreases — and barely: two scalars cannot grow a ring. The mechanism is demonstrated; competence is not claimed.)
This tiny example is deliberately restricted.
Two scalar parameters are nowhere near enough to learn rich morphogenesis.
But the important mechanism is already visible:
target behavior
↓
loss
↓
backpropagation through time
↓
local rule parameters
The global objective can train a local rule
This is the key idea.
The loss sees the whole final pattern.
But the update rule is applied locally and identically to every cell.
No cell receives coordinates saying:
you are the top-left corner of the target
The system must learn a local process whose repeated application causes the global structure to emerge.
That absence of positional instruction is what “self-organizing” means in this book: not a mystery fluid, but the concrete constraint that global order must arise from local shared updates alone.
That is why neural cellular automata are interesting.
They combine:
locality
weight sharing
recurrence
self-organization
learning
Designed local law vs learned local law
The inversion, stated once for the rest of Part V to inherit:
designed rule (Chapters 1–36)
specify: rule
observe: behavior
learning happens: nowhere (search selects; nothing trains)
learned rule (this chapter on)
specify: desired behavior (loss + target)
optimize: rule parameters (gradient descent)
learning happens: during training only —
deployment is frozen-parameter rollout
With it, the guardrails:
parameterized update
≠ intelligence
goal-conditioned behavior
≠ general problem solving
recovery after damage
≠ biological regeneration (task setup in Chapters 42–43,
not equivalence)
hidden channels (Chapter 39)
≠ meaning-bearing symbols unless demonstrated
successful training on one target
≠ general morphogenetic competence
What the full NCA adds
This chapter’s toy is a deliberately thin slice of Mordvintsev et al.’s Growing Neural Cellular Automata (2020). The full architecture, owned by the chapters ahead, adds:
this chapter
1 visible channel, cross-kernel perception,
2 scalar parameters, deterministic sync updates,
single target, fixed horizon
full NCA (Mordvintsev et al.)
16 channels: RGB + alpha "alive" + 12 hidden,
fixed Sobel-gradient perception (48-dim vector),
~8K-parameter residual update, zero-init do-nothing start,
stochastic per-cell update masks (no global clock),
alive masking (dead cells zeroed),
sample-pool training for persistence,
grow → persist → regenerate regimes by training setup
Nothing here contradicts that architecture; the toy demonstrates the differentiability mechanism while the chapters ahead supply the capacity, the stochasticity, and the training regimes. Conflating the two — assuming this demo already grows, persists, or regenerates — would be the first error of this part, so it is ruled out in advance.
The full per-cell update as a dataflow — every stage here is owned by a specific later chapter:
flowchart LR
S[state: 16 channels] --> P[perception: Sobel gradients]
P --> N[neural update: residual delta]
N --> M[stochastic mask: random subset fires]
M --> A[alive mask: zero the dead]
A --> S
Canonical toy-vs-full split, in one place:
| Aspect | This chapter’s toy | Canonical NCA (Mordvintsev et al.) |
|---|---|---|
| channels | 1 visible | 16: RGB + alpha + 12 hidden |
| perception | cross kernel | fixed Sobel gradients (48-dim) |
| update | 2 scalars | ~8K residual network, zero-init |
| scheduling | synchronous | stochastic per-cell masks |
| dead cells | none distinguished | alive masking at α > 0.1 |
| training | single target, fixed horizon | pool, persistence, damage regimes |
Differentiability changes what can be specified
With hand-written CA we specify:
rule
and observe:
behavior
With differentiable CA we can specify:
desired behavior
and optimize:
rule
That reverses the direction of the design problem.
But differentiable does not mean easy
Training through many recurrent steps creates familiar problems:
vanishing gradients
exploding gradients
unstable dynamics
short-horizon solutions
fragile attractors
A model may learn to produce the target at exactly step 32 and then immediately destroy it.
That is not persistent morphogenesis.
It is merely trajectory fitting.
We will solve these problems progressively rather than hiding them.
What we need next
A serious neural cellular automaton needs more expressive local computation than two trainable scalars.
The natural next step is:
neighborhood perception
↓
small neural network
↓
state update
The same tiny network is applied independently at every cell.
That gives us a learned local rule.
In the next chapter we will build exactly that.
Research
Mordvintsev, A., Randazzo, E., Niklasson, E. & Levin, M. — Growing Neural Cellular Automata: Differentiable Model of Morphogenesis (Distill, 2020). The primary source for this chapter and the rest of Part V: 16-channel state (RGB, alpha-alive, hidden), Sobel perception, residual update with zero-init, stochastic masks, alive masking, pool training, and the grow/persist/regenerate regime ladder. Everything the toy omits is specified here. https://doi.org/10.23915/distill.00023
Berto, F. & Tagliabue, J. — Cellular Automata (Stanford Encyclopedia of Philosophy). Situates the transition in one line the book relies on: CA built on neural networks with learned asynchronous rules for morphogenesis — the lineage this chapter joins, with the differentiability mechanism now demonstrated rather than cited. https://plato.stanford.edu/entries/cellular-automata/
Learned-rule vocabulary (Part V checkpoint)
| Concept | Meaning in this book |
|---|---|
| learned rule | update parameters optimized by training (this chapter) |
| cell state | full per-cell vector; here 1 visible channel, 16 with hidden from Ch39 |
| visible channels | rendered/output part of the state |
| hidden channels | internal working state, no predefined meaning (Ch39) |
| training | gradient updates to rule parameters |
| rollout | fixed-parameter execution; deployment is frozen rollout |
| regeneration | observed behavior under a damage task (Ch42–43), never biological equivalence |
| self-organizing | global order from local shared updates without positional instruction |
| organism (in these chapters) | shorthand for the grown target pattern; descriptive, as in Chapter 31, never a biological claim |