← Cellular Automata From First Principles

Make Experiments Reproducible

By this stage the book contains many systems whose behavior depends on parameters, seeds and implementation choices.

If we cannot reconstruct a run, we cannot really compare it.


Treat configuration as data

from dataclasses import dataclass, asdict


@dataclass(frozen=True)
class ExperimentConfig:
    model: str
    width: int
    height: int
    steps: int
    seed: int
    backend: str
    dtype: str
    params: dict = None

Then extend with model-specific parameters rather than hiding them in notebook cells.


Seed every relevant source of randomness

import random
import numpy as np
import torch


def seed_everything(seed):
    random.seed(seed)
    np.random.seed(seed)
    torch.manual_seed(seed)
    if torch.cuda.is_available():
        torch.cuda.manual_seed_all(seed)

Determinism can still depend on backend and operation, but explicit seeding is the minimum contract. The documented limits are worth stating plainly: identical seeds do not guarantee identical results across PyTorch releases, platforms, or CPU-versus-GPU execution — and some CUDA operations (convolution benchmarking, nondeterministic backward passes such as the one behind padded convolutions) need explicit flags beyond seeding. Seeding is necessary, not sufficient.


Save the exact configuration

import json
from pathlib import Path


def save_config(config, path):
    Path(path).write_text(
        json.dumps(asdict(config), indent=2),
        encoding="utf-8",
    )

A figure without the parameters that generated it is a dead artifact.


Record software context

Useful experiment metadata includes:

Git commit
Python version
library versions
device
operating system
random seed
input data hash

This does not guarantee identical results forever.

It makes discrepancies explainable.


Give each run an identity

A simple run directory might look like:

runs/
  2026-08-10T205700_rule110_seed42/
      config.json
      metrics.json
      final.npy
      frames/
      benchmark.json

Better still, derive a stable ID from normalized configuration plus code version. The full reproducibility chain, with each link’s failure mode:

    flowchart LR
    C[config + seed + version] --> R[run]
    R --> A[artifacts: raw + metrics]
    A --> P[replay from config]
    P --> K{bit-identical?}
    K -->|yes| T[trustworthy result]
    K -->|no| D[debug: backend, version, or nondeterminism]
  
InputWhy neededWhat can still vary
deterministic seedselects one realizationnothing, if backend is deterministic
backend + deviceexecution semantics differCPU vs GPU numerics
library versionsalgorithms change across releasescuDNN benchmarking choices
saved configreconstructs the runcode changes since (pin the commit)
raw artifactsre-derive figures without rerunstorage cost, format rot

Separate inputs, outputs and derived artifacts

inputs     → seeds, maps, targets
outputs    → raw states, checkpoints, metrics
derived    → plots, animations, summary tables

That distinction matters because a graph should be regenerable from raw results.


Replay should be a first-class operation

Replay is not an aspiration here — it is a function, shown complete for a Life configuration and verified bit-identical across repeated runs:

def replay(config):
    seed_everything(config.seed)
    rng = np.random.default_rng(config.seed)
    state = (rng.random((config.height, config.width)) < 0.4).astype(
        np.uint8
    )

    for _ in range(config.steps):
        state = life_step_np(state)

    return state

(life_step_np is Chapter 51’s NumPy life_step; the name marks it as the NumPy version, since Chapter 52 defines a tensor one.)

The strongest test is simple:

Can another process reconstruct this run from the saved configuration?

(Verified: two replays of a 48×48, 100-step configuration produce byte-identical final states, and the config serializes to JSON — reconstructability demonstrated, not asserted.)


Reproducibility enables debugging

When an interesting organism disappears after a code change, compare:

same config
old commit
new commit

Now the change in behavior is evidence about the implementation.

Without replay, it is merely a memory that something once looked different.


Search produces thousands of candidates.

A candidate record should contain enough information to recreate it:

candidate = {
    "score": 0.812,
    "seed": 3917,
    "parameters": params,
    "steps": 800,
    "model": "lenia",
}

Never save only the top PNG.


The experiment is the unit of knowledge

A useful mental model is:

claim
  ↓
configuration
  ↓
execution
  ↓
raw result
  ↓
metric / figure

That chain is what lets the book move from demonstrations toward real computational experiments.

Next we will use this structure to run parameter sweeps and benchmarks systematically.


Research

  • The Turing Way — Guide for Reproducible Research. The community handbook this chapter’s practices implement: configuration capture, run identity, inputs/outputs/derived separation, and the claim→configuration→execution→result chain as the unit of knowledge. Read it before inventing a lab workflow from scratch. https://book.the-turing-way.org/reproducible-research/reproducible-research/

  • PyTorch documentation: Reproducibility notes. The exact limits of seed_everything-style discipline: what seeding controls, where cuDNN benchmarking and nondeterministic algorithms break it, which flags restore determinism, and why CPU/GPU parity is not promised. Required reading before claiming any torch result reproduces. https://docs.pytorch.org/docs/stable/notes/randomness.html