Skip to content

Testing your model

Henad's testing kit checks a model's declarations against what its state actually does. A project made from the template already runs it over every model in models(), and a model added there is checked by the next cargo test.

Henad 0.3

This page describes Henad 0.3.

The template's test

The test sits at the foot of the template's src/lib.rs:

src/lib.rs
//! Models of the my-model project.

// Proving a type that holds wgpu handles `Send` or `Sync` walks wgpu-core's registries, deeper than the default
// limit of 128.
#![recursion_limit = "256"]

mod gpu_vote;
mod vote;

henad::include_shaders!();

use henad::authoring::register_gpu_grid_model;
use henad::authoring::register_grid_model;

/// Returns every model this crate provides.
///
/// # Errors
///
/// Returns [`henad::ModelSetError`] when two models share an id or an id breaks the id grammar.
pub fn models() -> Result<henad::ModelSet, henad::ModelSetError> {
    let mut models = henad::ModelSet::new(henad::build_info!());
    models.insert(register_grid_model::<vote::Vote>())?;
    models.insert(register_gpu_grid_model::<gpu_vote::GpuVote>())?;
    Ok(models)
}

#[cfg(test)]
mod tests {
    use henad::testing::{CheckSettings, TestDeviceRequest, assert_set_conforms, headless_test_device};

    #[test]
    fn every_model_conforms() {
        let models = super::models().expect("model ids are unique");
        let mut settings = CheckSettings::default();
        if let Some(device) = headless_test_device(&TestDeviceRequest::baseline()) {
            settings = settings.gpu(device);
        }
        assert_set_conforms(&models, &settings);
    }
}

The kit is henad::testing, behind the facade's testing feature, and the template turns the feature on for its tests alone:

Cargo.toml
[dev-dependencies]
henad = { version = "0.3", features = ["testing"] }

assert_set_conforms runs every check that applies to each model, and panics with every failure of the set. A set that passes prints which checks each model skipped, and why. cargo test hides that list for a passing test, and cargo test -- --nocapture shows it.

check_model_set returns the same report without asserting anything, and check_model checks one entry. Henad's example models use the report to check each model's skipped checks as well:

/// Returns the kit's settings over the example models, with a baseline device when this machine has an adapter, and
/// whether the settings hold a device.
///
/// The device requests `Limits::default()`, so a GPU model that only fits a raised limit fails to build here. Every
/// example model is meant to run on a stock WebGPU device.
fn kit_settings() -> (CheckSettings, bool) {
    let settings = CheckSettings::default();
    if let Some(device) = headless_test_device(&TestDeviceRequest::baseline()) {
        (settings.gpu(device), true)
    } else {
        log::warn!("checking the example models without their GPU checks: no adapter");
        (settings, false)
    }
}

/// Checks every example model once, asserting that each model passes and skips only what its backend, its declared
/// replay or a missing device rules out.
#[test]
fn the_example_models_conform() {
    let (settings, has_device) = kit_settings();
    install_panic_hook();
    let models = crate::example_models();
    let report = check_model_set(&models, &settings);
    report.assert_passed();
    for (entry, model_report) in models.iter().zip(report.reports()) {
        assert_skips_only_what_it_declares(entry, model_report, has_device);
    }
}

assert_skips_only_what_it_declares, beside the test, asserts that a model skips only the checks that its backend, its declared replay or a missing device rules out.

A device for the GPU checks

headless_test_device acquires a GPU device without a window, or returns None on a machine without a GPU. The GPU checks are then skipped and listed, and the CPU checks still run.

TestDeviceRequest::baseline() requests the limits that a browser offers by default, and the template's test uses it. A model that binds more storage buffers than the baseline allows is tested with TestDeviceRequest::raised(models.gpu_needs()), which requests the needs of the set's GPU models, and a browser at the baseline rejects it. Such a model also exempts DefaultsFit, which checks against the baseline on any device, as Settings and exemptions shows. features adds wgpu features to either request, and a test uses the henad::gpu::wgpu re-export for them, without its own wgpu dependency:

use henad::gpu::wgpu;
use henad::testing::{TestDeviceRequest, headless_test_device};

let request = TestDeviceRequest::baseline().features(wgpu::Features::TIMESTAMP_QUERY);
let Some(device) = headless_test_device(&request) else { return };

A missing feature returns None, as a missing device does, and returns it even under HENAD_REQUIRE_GPU. A test that requests an optional feature checks for None itself. headless_test_device returns None for either request on an adapter below the baseline.

Each test acquires its own device, as the template's test does. Clones of a device share one record of its errors, and a check treats any error it finds there as its own. An error that another test leaves on a shared device is then dropped, or reported as the failure of a check it has nothing to do with.

HENAD_REQUIRE_GPU=1 turns a missing device into a failure, and fails every check a missing device would skip. Set it wherever a GPU must be there.

The checks

Check Pins
ModelId, ParamIds, StatLabels, ActionIds The id meets the grammar, no parameter, stat or action id is declared twice, and the command line can refer to every parameter and action id
Palette A declared palette has colours
Metadata The backend, the structure, the topology hint and the device demand agree
DefaultSetup The declared defaults pass RunSetup::from_parts, as the app's Build checks them
DefaultsFit A GPU model's defaults fit a stock WebGPU device, checked without building
ApplyModes A live parameter is accepted and a reload parameter is rejected, exactly as declared
Views The factory returns the declared backend, and its grid, point or edge views match the topology hint
ParallelJobs Only a CPU model reports how many jobs a step splits into
Actions Every declared action is accepted and an index past the last is rejected, on a GPU model at its declared defaults
StatCount stats returns a value for every entry of STATS
ThreadCount A CPU model's stats and exported state are the same at 1 and 7 threads
SameSeed, SeedSensitivity Two runs on one seed agree, and two seeds differ in some stat or in the exported state
SamplingCadence A run sampled every tick ends where a run sampled every seventh tick ends
BaselineBuild A GPU model builds on the device at its declared defaults, and the capacity check agrees that it fits
FullSubmission One submission of 64 steps reads back what 64 submissions of one step do, the OS watchdog trap of the GPU backend. A model that does not replay exactly and declares stats reads back some stat that is not zero
SampledSlice A sampled slice of steps reads back what a snapshot does

A check that builds the model sets the grid, population and world sizes small, so a model whose defaults hold ten million agents needs no settings. ThreadCount sets the size itself, for a step to split into 14 jobs. At one job a kernel that shares state between chunks agrees with itself at any thread count. The GPU checks build at the declared defaults, the size the app builds first.

A GPU model skips ThreadCount, since a pool width never reaches its kernels. A model that declares REPLAYS_EXACTLY = false skips SameSeed, SeedSensitivity and SamplingCadence.

Settings and exemptions

CheckSettings holds what the checks share. ticks sets how many ticks a check steps, thread_counts sets the two pool widths that ThreadCount compares, and set_text sets a parameter in every check that builds the model, as --set reads it. An override of a parameter the model does not declare, or a value that the parameter rejects, fails every check that builds the model, on a machine without a device as well.

A check that the model cannot meet for an honest reason gets an exemption, recorded in the model's report:

let exempt = CheckSettings::default().exempt(model.id(), ModelCheck::SeedSensitivity, "draws no random number");

A model whose defaults need a device raised past the baseline fails DefaultsFit, and exempts it with its reason. An exemption of a check that does not apply to the model fails that check. The report of a set lists every model that the settings refer to and that is missing from the set, and a renamed model or parameter cannot leave a stale setting behind.

check_model returns the report instead of panicking, and its caller installs the panic hook first with henad::install_panic_hook. Otherwise the failure from a kernel panic shows no file:line.

Where the GPU checks run

The GPU checks run wherever headless_test_device finds a device: on your own machine, and on CI that installs a driver. A runner on GitHub has no GPU. The template's workflow installs lavapipe, a Vulkan driver that runs on the CPU, with scripts/install-lavapipe.sh, and sets HENAD_REQUIRE_GPU=1.

Lavapipe catches zero readbacks and a wrong step count, and has no watchdog. FullSubmission guards against the OS watchdog, which fires only on real hardware. There, too many passes in one command buffer leave every later readback reading zero, with no error. Run cargo test on your own machine's GPU before you trust a GPU model.

What stays hand-written

The kit checks that a model keeps its contract. It has no oracle for what the model should compute, and the tests that check the rule itself stay yours:

  • A pattern drawn by hand. The Game of Life tutorial checks a blinker against its known next states.
  • A closed form or an invariant. Checking the rule itself lists the kinds Henad's example models use.
  • A busier thread-count test. ThreadCount runs a small configuration. A model that draws random numbers in several places can carry its own thread-count test at a size where every path runs.
  • A comparison with another engine, from a written procedure.
  • A GPU port against its CPU model. A port that seeds itself through its CPU model's init starts on the same state at tick 0, and a test can compare the two backends as long as the model itself has no randomness.

Next