Build a Cellular Automata Laboratory
We now have enough pieces to stop thinking in terms of isolated scripts.
A useful cellular-automata laboratory should let us define, run, measure, replay and visualize experiments across many kinds of automata without hiding the mechanisms we spent the whole book learning.
The laboratory is not the model
Keep these concerns separate:
model definition
execution backend
experiment configuration
metrics
artifacts
analysis
A rule should not know where its PNG is saved.
A renderer should not decide how a cell updates.
One possible project structure
ca_lab/
├── models/
│ ├── elementary.py
│ ├── life.py
│ ├── lenia.py
│ └── nca.py
├── backends/
│ ├── numpy_backend.py
│ ├── torch_backend.py
│ └── fft_backend.py
├── experiments/
│ ├── config.py
│ ├── runner.py
│ └── sweep.py
├── metrics/
│ ├── activity.py
│ ├── entropy.py
│ ├── recurrence.py
│ └── robustness.py
├── render/
│ ├── figures.py
│ └── animation.py
├── artifacts/
├── tests/
└── cli.py
The exact folders are less important than the boundaries.
A model contract
class Model:
def initial_state(self, config):
raise NotImplementedError
def step(self, state):
raise NotImplementedError
A learned model can additionally expose parameters or checkpoints.
A deterministic classical rule does not need to pretend it has them.
An experiment contract
@dataclass
class Experiment:
config: ExperimentConfig
model: Model
observers: list
def run(self):
state = self.model.initial_state(self.config)
for step in range(self.config.steps):
state = self.model.step(state)
for observer in self.observers:
observer(step, state)
return state
This creates one execution path for:
Rule 30
Life
traffic
forest fire
Lenia
NCA
without forcing their internal state to be identical.
A command-line interface makes experiments concrete
Conceptually:
ca-lab run experiments/rule110.yaml
ca-lab sweep experiments/lenia-search.yaml
ca-lab replay runs/a81f2c
ca-lab render runs/a81f2c
ca-lab benchmark benchmarks/neighborhoods.yaml
A CLI is useful because it turns an experiment into something that can be executed outside a notebook.
Keep notebooks at the edge
Notebooks are excellent for:
exploration
explanation
plotting
interactive inspection
They are poor as the only location of core model logic.
Prefer:
library code
↓
experiment runner
↓
notebook imports results
rather than copying the implementation between notebooks.
Test the mechanisms
The laboratory should contain small invariant tests — real ones, reusing owned machinery. Chapter 11’s traffic system supplies the example:
def test_rule184_conserves_cars():
state = make_road(density=0.35, seed=42)
assert state.sum() == traffic_step(state).sum()
(Verified passing: the invariant from the traffic chapter becomes a regression test here, which is exactly how model contracts earn their keep.)
and integration tests:
saved config can replay
artifact manifest points to existing result
NumPy and PyTorch reference implementations agree within tolerance
Tests make the experimental infrastructure trustworthy — with the earned principle stated plainly:
Surprising output should trigger verification before interpretation.
This book’s own history supports it: undefined helpers, wrong signatures, key mismatches, FFT miscentering, padding-semantics drift, and uncalibrated thresholds were all found by executing the code. Every one of them once produced a plausible-looking artifact.
test passes
≠ model scientifically validated
reference match
≠ real-world validity
plausible image
≠ correct implementation
Store raw evidence before summaries
A good run might produce:
config.json
metadata.json
metrics.csv
final_state.npy
checkpoints/
frames/
Then derived outputs:
summary.json
plots/
animation.mp4
report.md
If a plot is wrong, regenerate it from raw output.
Do not rerun an expensive experiment merely because a label was misspelled.
Make comparison a first-class operation
ca-lab compare runs/a81f2c runs/b114e9
A comparison might show:
parameter differences
metric differences
runtime differences
final-state distance
behavioral fingerprints
Now changing an implementation or rule becomes an explicit experiment. The laboratory’s full loop — hypothesis in, compared evidence out:
flowchart LR
H[hypothesis] --> C[config + seed + version]
C --> R[runs: raw artifacts]
R --> V[validate: invariants + replay]
V --> M[derive metrics + figures]
M --> K[compare runs]
K --> N[conclusion or new hypothesis]
Lab operations compared by what they consume and produce:
| Operation | Consumes | Produces | Evidence boundary |
|---|---|---|---|
| run | config + model | raw states, metrics | single realization only |
| sweep | parameter grid | records table | coverage, not proof |
| replay | saved config | identical (or revealingly different) states | tests reproducibility, not truth |
| render | raw outputs | figures, animations | presentation, never new data |
| compare | two run IDs | differences + fingerprints | difference observed, cause still open |
The laboratory preserves what the book teaches
The architecture should not abstract away the subject until all automata look like generic black boxes.
The point is the opposite.
It should preserve inspectable concepts:
state
neighborhood
rule
update schedule
measurement
search
perturbation
while removing accidental duplication around them.
We are ready for the capstone
A laboratory becomes useful when it can challenge a result we actually want to believe. For the capstone we return to the richest hand-designed system in the book, Lenia, and use the machinery of measurement, search, reproducibility and rejection to test a candidate rather than admire it.
The next and final chapter will use the whole workflow:
define a search space
run reproducible experiments
find a candidate
measure it
perturb it
compare alternatives
inspect its computation
produce publication artifacts
state only what the evidence supports
That is the complete journey of the book.
Research
Python documentation:
dataclasses— Data Classes. The contract machinery reused here: frozen configuration objects, generated initializers, and explicit field defaults — the documented basis for theExperiment/ExperimentConfigpattern from the previous two chapters. https://docs.python.org/3/library/dataclasses.htmlThe Turing Way — Guide for Reproducible Research. The lab-level practices this chapter assembles: run identity, raw-evidence-before-summary, comparison as an operation, and the claim→configuration→execution→result chain as the unit of knowledge. https://book.the-turing-way.org/reproducible-research/reproducible-research/