Skip to content

Determinism and testing

Henad runs every kernel in parallel and gives no guarantee about which chunk runs on which core. The answer still has to come out the same every time, and this page covers the rules that make sure it does.

The engine seeds a chunk's RNG from the tick's seed and the chunk's index, and from nothing a worker mutates, which keeps a result independent of how rayon schedules the chunks. The same has to hold on the web, where the pool width is whatever navigator.hardwareConcurrency reported.

Reductions obey the same rule. reduce_chunks folds its partials in chunk order rather than completion order, a Tally merges in chunk order, and the scatter totals in fixed point before touching f32. Float addition is not associative, so if any of these folded in arrival order instead, a stat would depend on the machine that ran it.

Every draw needs its own next_bits

Feeding one word to two draws correlates them, and neither the compiler nor a test will tell you it happened. The split into an advancing call plus pure draws exists so that the draws can be parity-tested against WGSL, and sharing one word between draws defeats the point of that split.

The thread-count test

The ThreadCount check of the testing kit runs this comparison on every CPU model, at a size that splits a step into 14 jobs. Both agent models also carry their own results_do_not_depend_on_the_thread_count test, with a busier configuration than the kit uses. If your model draws random numbers during a step, consider adding that test too. The test builds its own thread pools with rayon. In a project made from the template, add rayon = "1" under [dev-dependencies] first. Cargo resolves it to the same rayon that Henad uses.

#[test]
fn results_do_not_depend_on_the_thread_count() {
    fn run(threads: usize) -> Vec<u32> {
        let pool = rayon::ThreadPoolBuilder::new()
            .num_threads(threads)
            .build()
            .expect("rayon pool");
        pool.install(|| {
            let params = vec![ParamValue::U32(20_000), ParamValue::F32(800.0), ParamValue::F32(800.0)];
            let mut state = State::from_params(&params);
            for _ in 0..30 {
                state.step();
            }
            let lanes = state.lanes();
            lanes
                .pos_x
                .iter()
                .zip(&lanes.pos_y)
                .flat_map(|(x, y)| [x.to_bits(), y.to_bits()])
                .collect()
        })
    }
    assert_eq!(run(1), run(7), "boid positions depend on the thread count");
}

Three details of the test matter.

  • Compare bits rather than floats. Two NaNs with the same bits compare equal, where == fails, and -0.0 and 0.0 differ, where == calls them equal. An approximate comparison would hide a last-bit difference altogether.
  • Use a population that spans several chunks, preferably one that is not a multiple of CHUNK, so that the ragged final chunk is covered too.
  • Use thread counts that are not multiples of each other, since running 1 against 7 splits the work completely differently.

To run a single test by name:

cargo test results_do_not_depend_on_the_thread_count

In a clone of Henad's repository, -p henad-models runs the example models' tests of that name.

Network models

After init, a network model draws random numbers in three places, and each has its own stream.

The node pass is seeded like an agent pass. The engine passes it a per-tick seed, and run_pass splits that seed per chunk. The global pass runs on one thread and draws from a single stream, in the order that the model requests numbers. That stream carries over from one tick to the next. An action draws from a third stream, seeded separately from the other two streams. A press advances neither the node pass's seed nor the global pass's stream.

Node positions are outside the contract. The layout runs when a snapshot is published, under a time budget, and never inside step. How far it has moved the nodes by a given tick depends on how many publishes there were and how many iterations each publish fitted in. Neither count is a function of the tick. The layout itself is deterministic for a fixed number of iterations, and layout.rs carries its own thread-count test. Keep positions out of anything a tick decides. Team Assembly reads a position in its global pass to place newcomers next to their team's first incumbent, and that placement only changes the picture.

The order of neighbours inside a row is deterministic. It follows from the sequence of edits made to the graph and from nothing else, and a repack keeps every row in order. The order carries no meaning, though. Removing an edge moves the last entry of each affected row into the gap, and flipping the direction rebuilds every row in edge-list order. A kernel can walk a row in order, but it must not give an entry a meaning by where it sits, for example by reading the first entry as the oldest edge.

Each network model carries two tests. If your model draws random numbers or writes anything in prepare_view, write both tests for it too.

results_do_not_depend_on_the_thread_count runs a busy configuration at 1 and at 7 threads, calls prepare_view after every tick, and compares the two runs bit for bit. The population spans several chunks, and a parameter changes halfway through. Virus on a Network runs with keep_rewiring on and switches directed on at the midpoint. It compares the state and timer lanes, the edge list and the edge colours. Team Assembly lowers max_downtime at the midpoint. It compares the occupied slots, the spawn and team ticks, the positions, the edge list with its colours, the node colours and the bits of every stat.

results_do_not_depend_on_the_publish_cadence runs the same configuration twice, calling prepare_view after every tick in one run and never in the other, and asserts that both runs end in the same state. The second run then calls prepare_view once, and the colours it paints have to match the first run's colours. prepare_view can write to the graph and the lanes, and the test pins that nothing it writes feeds back into a tick.

Virus on a Network leaves positions out of both comparisons. Team Assembly keeps them in, since its own tick places newcomers and neither test runs the layout.

The testing kit

The testing kit runs the thread-count, seed and sampling comparisons on every model of a set, beside the checks on each model's declarations. A project made from the template runs it from cargo test.

Checking the rule itself

Determinism aside, the model's rule still needs its own oracle. The repository leans on three kinds, in descending order of strength.

A bit-identical reference. This is only available when nothing in the model is stochastic. gpu_game_of_life is checked against its CPU counterpart this way, and Game of Life is checked against fixtures recorded from NetLogo.

A closed form. SIR's infection rate is checked against 1 - (1 - beta)^n over a large grid, with a tolerance band derived from the binomial spread rather than picked by hand, and its recovery rate is checked the same way.

An invariant. This is the weakest kind of oracle, but it is always available. Examples in the repository include population being conserved, transitions only going forwards, deliveries never decreasing, ants staying on the lattice, inside the world and off obstacles, and a cell with no ant on it decaying by exactly the evaporation rate.

The consistency tests live in crates/henad-models/tests/ and are named consistency_<model>.rs.

Consistency fixtures

A fixture recording another engine's output has to come from a written procedure, and a generation script does not qualify. The procedure goes in crates/henad-models/tests/fixtures/docs/ for a human to run.

A driver script would presume the reference engine is installed, which no future collaborator can be expected to have. The committed fixture together with its procedure is the reproducibility record.

Where the reference engine is code rather than a GUI, a small committed program is the procedure, and that is fine.

Never generate a fixture from Henad

A fixture produced by the engine under test only proves that the engine agrees with itself, so a test against it passes forever and means nothing.

For a stochastic model the two engines draw from different generators, and a fixture cannot then be compared point by point. scripts/compare_sir.py instead compares the distribution of summary statistics over many replicates, against margins derived from Henad's own measured run-to-run spread.

The two network models are compared with NetLogo the same way. Their procedures are virus_network_fixture.md and team_assembly_fixture.md in crates/henad-models/tests/fixtures/docs/. scripts/compare_network.py judges their runs the same way compare_sir.py judges SIR runs. On the Henad side, crates/henad-models/tests/consistency_virus_network.rs and consistency_team_assembly.rs hold the consistency tests, and they need no reference engine. The Virus on a Network tests check every rate against the edge list instead of the model's own rows, and a row that drifted from the list fails them too. The Team Assembly tests compare the lanes and the graph with a scan or a closed form instead of the model's own bookkeeping.

Before calling it green

In a project made from the template, this script runs the stages its CI runs:

scripts/ci.sh

GPU tests skip silently on a machine with no adapter, so set the environment variable that turns the skip into a failure:

HENAD_REQUIRE_GPU=1 cargo test

In a clone of Henad's repository, ./check.sh runs Henad's own check set, and HENAD_REQUIRE_GPU=1 cargo test --workspace --all-targets runs its tests.

Next