Library M10a: the testing kit¶
M10a is the first part of the tenth of the eleven milestones of #48, which turns Henad into a library published on crates.io.
henad_explore::testing, behind henad-explore's newtestingfeature, checks any model entry against what its state does and returns a report instead of panicking. It holds twentyModelChecks in four groups,CheckSettings,ModelReport,TestDeviceRequestand the native-onlyheadless_test_device. The example models' registry tests became one kit run that asserts both conformance and each model's skipped checks, and every test device outside henad-compute now comes fromheadless_test_device. A review pass followed. It reorderedFullSubmissionso a poisoned device cannot pass it, made the settings' mistakes visible, and gave every check a self-test.cpu/grid_engine.rs's thread-count test takes a 128 by 1024 grid, the named exception of [4.9.4].
State before¶
48-library stood at 4e1bd2b (M8), with M9's work uncommitted in the main checkout, HENAD_REQUIRE_GPU=1 ./check.sh at 1046 tests.
The maintainer committed M9 as d374503 before the review pass.
crates/henad-models/src/tests/registry.rs held 22 tests, 14 of them contract tests over example_models() that picked GPU entries by backend.
They protected the ten example models alone, and a downstream author had none of them.
Headless device set-up and HENAD_REQUIRE_GPU parsing were copied six times: henad-compute's and henad-models' support.rs, henad-explore's headless_device and baseline_device, the CLI's SmallGpuGrid and golden test, and the app's AppState::headless.
cpu/grid_engine.rs's results_do_not_depend_on_the_thread_count ran a 64 by 64 grid, one rayon leaf, and compared one job with itself.
What was done¶
The kit¶
crates/henad-explore/src/testing/ is compiled with the testing feature, and always for henad-explore's own tests.
mod.rsholdsModelCheck(20 variants,ModelCheck::ALLas a slice in run order),check_model,check_model_setandassert_set_conforms. Each check runs insidecatching, orcatching_onfor a GPU model. A GPU check then waits for the device, and a fault its work left in the sink is that check's failure, so no fault outlives the check that raised it. A check never panics on a model's behalf. Stats are compared bit for bit throughstat_bits, histograms included.check_modelfails a check thatCheckSettings::exemptexempts when the check does not apply (another backend, or a comparison of runs on a model that does not replay exactly). UnderHENAD_REQUIRE_GPU, a check a missing device would skip fails instead.check_model_setreturns aSetReport, the reports in the set's order plusunknown_models, every model id an override or an exemption names and the set lacks.assert_set_conformsinstalls the panic hook, panics with the report when anything failed, and prints it when the set passes, skipped checks included. libtest shows the print under--nocapture.report.rsholdsModelReport(model_id,failures,skipped,passed,assert_passed, plusskip_reasonandthread_count_jobs),SetReport,CheckFailure,SkippedCheckandSkipReason:OtherBackend,InexactReplay,Exempt(reason),NoDevice,NativeOnlyandOneJob. A report's first line is a summary, "Model 'x' failed 2 of its checks: Actions, Views.", and the failures and the skipped checks follow, one to a line.settings.rsholdsCheckSettingswithgpu,ticks(20 by default, at leastMIN_TICKS, 8),thread_counts(1 and 7),set_textandexempt. A check that builds takesgrid_width,grid_height,world_widthandworld_heightat 128 andnum_agentsat 256, never above the default and clamped to the bounds, then the overrides.declared.rs:ModelId(throughModelSet::insertinto an empty set, so the grammar has one home),ParamIds(naming a repeat of an engine-prepended id),StatLabels,ActionIds,Palette,Metadata,DefaultSetupandDefaultsFit.built.rs:ApplyModes,Views(the factory's arm, then the CPU views or the GPU layers against the hint),ParallelJobs,ActionsandStatCount, each on a build of its own. A GPU model'sActionsbuilds at the declared defaults.determinism.rs:ThreadCount,SameSeed,SeedSensitivityandSamplingCadence, throughRunSetup::from_partsandSimulation, each comparing the stats and, for a CPU model,write_state's bytes.ThreadCountsizes the work for twice the high thread count, 14 jobs: an agent or network model takesnum_agentsat 14 chunks fromStructure::AgentsorStructure::Network, and a grid model the largestgrid_heightthat splits into no more than 14 jobs, found by doubling and bisecting over builds, 896 rows at 128 columns. An override of the size parameter keeps its value, one job at every size within bounds skips the check asOneJob, and wasm32 skips it asNativeOnly. A model that does not replay exactly skipsSameSeed,SeedSensitivityandSamplingCadence.gpu.rs(native only):BaselineBuild(the build succeeds exactly when the shortfalls against the device's limits are empty),FullSubmissionandSampledSlice.FullSubmission, on a model that replays exactly, first runs 64 submissions of one step and reads their stats, then the old full submission (64 steps and the snapshot passes in one command buffer, then a blocking readback), and compares the two. A device the full submission poisons reads zeros in every later run, and the other order compares zeros with zeros. It fails on a lost device, a readback that did not land, or a tick other than 64.gpu_boidskeeps the old form after the full submission: some stat is not zero.device.rs:TestDeviceRequest(baseline,raised(needs),features),headless_test_device, which asks for a high-performance adapter, returnsNonefor a missing optional feature even underHENAD_REQUIRE_GPU, and attaches the adapter'sRuntimeInfo, and the one reading ofHENAD_REQUIRE_GPUthe kit makes.henad_compute::gpu::wgpure-exports wgpu, the path [4.1] names, andgpu_game_of_life's timing test nameshenad_compute::gpu::wgpu::Features::TIMESTAMP_QUERYthrough it.
The old registry tests, check by check¶
No assertion of the 14 contract tests was dropped.
| Old test | Covered by |
|---|---|
every_default_setup_passes_the_checks_of_from_parts (the M6 review's check) |
DefaultSetup, both assertions. [4.8]'s DefaultsFit reads the GPU limits only and does not cover it |
declared_apply_mode_matches_what_the_state_accepts |
ApplyModes |
declared_topology_matches_the_views_the_state_returns |
Views, CPU arm |
declared_topology_matches_the_layers_a_gpu_state_publishes |
Views, GPU arm |
only_a_cpu_model_reports_how_far_a_step_splits |
ParallelJobs |
declared_metadata_matches_the_entry_it_describes |
Metadata (capacity, needs, CPU replay, structure against hint and backend), and Views for the factory's arm |
every_declared_action_is_accepted_by_the_state |
Actions, at the defaults on a GPU model as before |
action_ids_are_unique_within_a_model |
ActionIds |
a_declared_palette_has_colours_in_it |
Palette |
every_declared_stat_series_gets_a_value |
StatCount |
every_gpu_model_builds_on_a_baseline_device |
BaselineBuild (the build and the agreeing demand) and DefaultsFit (the shortfalls against Limits::default()) |
every_gpu_entry_reports_its_capacity |
Metadata (a GPU demand of some bytes, no CPU demand) |
a_full_submission_executes_every_step |
FullSubmission. For a model that replays exactly, the non-zero stat became a comparison with single steps run before it |
a_sampled_slice_reads_back_what_a_snapshot_does |
SampledSlice, and the fault-sink check of every GPU check |
Kept beside example_models(): every_gpu_entry_needs_the_bindings_its_widest_pass_binds, a_model_too_large_for_the_device_is_reported, every_example_model_joins_the_set, an_example_entry_inserted_into_another_set_keeps_its_source and a_gpu_entry_refuses_to_build_without_a_device.
[4.8] drops StorageBuffers from the kit, and the first of these keeps that assertion for the example models.
The engine's own tests stay: a_model_that_panics_while_building_comes_back_as_a_fault, a_kernel_that_panics_mid_step_keeps_its_location and a_model_entry_can_be_shared_between_threads.
the_example_models_conform runs the kit once through check_model_set, asserts that the set report passed, then checks every model's skip list exactly: the four GPU checks as OtherBackend on a CPU model, ThreadCount as OtherBackend on every GPU model, SameSeed, SeedSensitivity and SamplingCadence as InexactReplay on gpu_boids, and NoDevice on the GPU models' builds when the machine gives no device.
It also asserts that every CPU model reached more than one job.
the_example_gpu_models_are_the_gpu_backend_entries asserts that the GPU models are the four gpu_ entries, picked by backend, and runs no kit.
Test devices¶
- henad-models'
src/tests/support.rsis deleted. Each GPU model'sheadless_context, the tutorial's GPU tests, the parity tests andsimulation.rscallheadless_test_device(&TestDeviceRequest::baseline()). gpu_game_of_life's timing test asks forTestDeviceRequest::baseline().features(henad_compute::gpu::wgpu::Features::TIMESTAMP_QUERY), and still skips on an adapter without it.- henad-explore's
headless_deviceandbaseline_deviceare one line each over the kit, raised to the example set's needs and at the baseline. - The CLI's
SmallGpuGrid::new, its golden test'shas_adapterand the app'sAppState::headlessgo throughTestDeviceRequest::raised. - henad-compute keeps its own helper, as [4.8] says.
- henad-models takes henad-explore as a path dev-dependency with
testing, and its unusedpollsterdev-dependency went. henad-cli and henad-app addtestingto the henad-explore they already depend on, for their tests.
Kit self-tests¶
crates/henad-explore/src/tests/kit.rs runs the kit over each model of tests/broken.rs and expects exactly the failing checks. Fourteen of the twenty checks have a self-test that fails them, and the issues below list the other six.
| Self-test | Model or wrapper | Fails exactly |
|---|---|---|
a_bad_id_fails_the_model_id_check |
BadId ("Bad Id") |
ModelId |
a_parameter_repeating_an_engine_id_fails_the_param_ids_check |
DeclaresNumAgents, an agent model |
ParamIds |
a_repeated_stat_label_fails_the_stat_labels_check |
RepeatsStatLabel |
StatLabels |
a_repeated_action_id_fails_the_action_ids_check |
RepeatsActionId |
ActionIds |
an_empty_palette_fails_the_palette_check |
EmptyPalette |
Palette |
a_state_that_accepts_every_edit_fails_the_apply_modes_check |
BuggyState over sir, Bug::AcceptsEveryEdit |
ApplyModes |
a_state_that_hides_its_grid_fails_the_views_check |
Bug::HidesGrid |
Views |
a_state_that_refuses_its_actions_fails_the_actions_check |
Bug::RefusesActions |
Actions |
a_state_that_drops_a_stat_fails_the_stat_count_check |
Bug::DropsLastStat |
StatCount |
a_build_that_reads_the_pool_width_fails_the_thread_count_check |
ReadsPoolWidth, as many live cells in its first row as the pool has threads |
ThreadCount |
a_shared_accumulator_fails_the_thread_count_check |
SharedAccumulator, a Mutex<u64> its chunks share |
ThreadCount among others, at 14 jobs |
a_build_that_differs_from_the_last_fails_every_check_comparing_two_builds |
CountsBuilds, a static counter mixed into init |
ThreadCount, SameSeed, SamplingCadence |
a_model_that_draws_no_random_number_fails_the_seed_sensitivity_check |
DividesByParam at its defaults |
SeedSensitivity, and passes with the exemption the report lists |
a_view_preparation_that_writes_a_lane_fails_the_sampling_cadence_check |
CountsViews, a network model whose prepare_view counts in a lane |
SamplingCadence |
a_full_submission_that_reads_zeros_fails_the_full_submission_check |
ZeroesFullSubmissions over gpu_game_of_life |
FullSubmission |
a_kernel_panic_fails_every_check_that_steps |
DividesByParam at divisor=0 |
the four determinism checks, naming divide by zero and broken.rs |
a_build_panic_fails_every_check_that_builds |
DividesByParam at init_divisor=0 |
the nine checks that build |
a_stat_that_stops_being_finite_breaks_no_contract |
InverseOfCountdown |
SeedSensitivity alone |
an_exemption_of_a_check_that_does_not_apply_fails_it |
Game of Life exempting FullSubmission |
FullSubmission |
DividesByParam and InverseOfCountdown draw no random number and fail SeedSensitivity on purpose. The grid models above come from a grid_model! macro in broken.rs, each with one broken declaration or build.
BuggyState replaces RefusesActions, which run_control.rs used, and forwards the views it did not.
ZeroesFullSubmissions shares one stopped flag among every state its entry builds, since every buffer of a device the watchdog stopped reads zero. Put back in the old order (the full submission first), FullSubmission passed that model. The mutation was run and reverted.
Three more: Game of Life passes at 14 jobs and skips the four GPU checks as OtherBackend, a set report names game_of_lfe and missing from an override and an exemption, and assert_set_conforms panics with "1 of 1 models failed their checks" and the model's summary line.
No self-test covers Metadata, DefaultSetup, DefaultsFit, ParallelJobs, BaselineBuild or SampledSlice. A model breaking one of them needs a hand-built entry, or a GPU model past the baseline.
Correction after a later review: register_grid_model accepts a default outside its bounds, so DefaultSetup needs no hand-built entry, and a_default_outside_its_bounds_fails_every_check_through_a_run_setup now covers it.
The review¶
The maintainer's review of the first pass came in eight parts, all folded in above.
FullSubmissioncompared zeros with zeros on a poisoned device. The single steps now run first, a lost device fails, andZeroesFullSubmissionspins it.- A missing device under
HENAD_REQUIRE_GPUskipped the GPU checks silently. They now fail, through the kit's one reading of the variable, andassert_set_conformsprints the skipped checks of a set that passes. SeedSensitivitycompared the stats alone. It now compares the exported state too, and skips asInexactReplayon a model that does not replay exactly.- Settings naming a model the set lacks were ignored.
SetReport::unknown_modelsnames them, and an exemption of a check that does not apply fails it. - A GPU model's
Actionsnow runs at its declared defaults, as the old test did. - Self-tests for every check that had none, the deterministic thread-count one included, and
kit.rs's doc and AGENTS.md no longer say "only that one". ModelCheck::ALLbecame a slice,check_modelandcheck_model_settake#[must_use],ThreadCountskips asNativeOnlyon wasm32, leftover faults are reported against the check that left them,CheckSettings::ticksrefuses fewer than 8, the registry runs the kit once,assert_passedopens with a summary line, and the determinism page says a model past the baseline exemptsDefaultsFit.- Rerun, remeasure and propose two commits, below.
The named exception¶
cpu/grid_engine.rs's results_do_not_depend_on_the_thread_count takes a 128 by 1024 grid, 16 jobs of 64 rows, with its assertion unchanged.
[4.9.4] asks for this in a commit of its own.
Docs and AGENTS.md¶
authoring/determinism.md's "Tests the registry brings" became "The testing kit", with the example models' test and an exemption included by--8<--region fromregistry.rsandkit.rs, a table of the checks, the sizes, the skips, theDefaultsFitexemption for a model past the baseline, the settings' mistakes and the panic hook.registering.md,porting.md,parameters.mdandstatistics.mdname the kit's check where they said "a registry test".- AGENTS.md: the test-module rule and the two
support.rsfiles, the dev-only edge from henad-models to henad-explore, the crate list, the registry tests, atesting/passage under henad-explore, the broken models andkit.rs, "Adding a new model", the thread-count and watchdog notes, andBaselineBuildin place ofevery_gpu_model_builds_on_a_baseline_device. - CHANGELOG: Added lines for the kit,
headless_test_devicewithTestDeviceRequest, andhenad_compute::gpu::wgpu. Nothing public was removed or changed, and no line carries "Breaking:".
Edited tree¶
.
├── AGENTS.md ~ the kit, the test devices, the dev-only edge
├── CHANGELOG.md ~ M10a's Added lines
├── Cargo.lock ~ henad-models' dev-dependencies
├── zensical.toml ~ nav entry #39
├── crates/henad-compute/src/
│ ├── gpu/mod.rs ~ pub use wgpu
│ └── cpu/grid_engine.rs ~ the thread-count test at 128 by 1024
├── crates/henad-explore/
│ ├── Cargo.toml ~ testing feature
│ └── src/
│ ├── lib.rs ~ pub mod testing
│ ├── testing/ + mod, report, settings, device, declared, built, determinism, gpu
│ └── tests/
│ ├── mod.rs ~ mod kit
│ ├── kit.rs + a self-test per check, the set report, assert_set_conforms
│ ├── broken.rs ~ grid_model!, the agent, network and wrapper fixtures, BuggyState
│ ├── run_control.rs ~ BuggyState in place of RefusesActions
│ └── support.rs ~ headless_device and baseline_device over the kit
├── crates/henad-models/
│ ├── Cargo.toml ~ henad-explore with testing, pollster out
│ └── src/
│ ├── gpu_{game_of_life,sir,boids,ants}/mod.rs ~ headless_test_device, the timing request
│ └── tests/
│ ├── mod.rs ~ support gone
│ ├── support.rs - replaced by headless_test_device
│ ├── registry.rs ~ one kit run with the skip lists, the guards
│ ├── simulation.rs ~ headless_test_device
│ └── tutorial/{gpu_life,gpu_foraging,parity}.rs ~ headless_test_device
├── crates/henad-cli/
│ ├── Cargo.toml ~ henad-explore with testing for the tests
│ ├── src/lib.rs ~ SmallGpuGrid over the kit
│ └── tests/golden.rs ~ has_adapter over the kit
├── crates/henad-app/
│ ├── Cargo.toml ~ henad-explore with testing for the tests
│ └── src/state.rs ~ AppState::headless over the kit
└── docs/
├── authoring/{determinism,registering,porting,parameters,statistics}.md ~ the kit
└── developing/agent-record/20261002-39-library-testing-kit.md +
State after¶
M10a is implemented with its review and uncommitted on 48-library, on top of d374503, in the main checkout as asked. Nothing is staged.
HENAD_REQUIRE_GPU=1 ./check.shpasses: 1057 tests, none failed, with the wasm32 typecheck, packaging, cargo-deny, docs and the web build. M9 passed at 1046, and the first M10a pass at 1042. The 14 contract tests became 2, andkit.rsholds 23.- check.sh's
--all-featureswasm32 line typechecks henad-explore withtesting. A first pass warned onTestDeviceRequest::needs, which only native code reads, and it is gated. uv run --locked zensical buildpasses with no issues, and the determinism page renders both included regions.cargo clippy --workspace --all-targets --all-features -- -D warnings -W clippy::allpasses.
The kit's wall time¶
Measured under cargo test on this machine, where the workspace members build at opt-level 0 and dependencies at 2, with HENAD_REQUIRE_GPU=1 and the baseline device.
- First pass,
the_example_models_conformalone with--exact --test-threads=1, as a whole process: 10.33 s on the first run, then 5.44 s and 5.44 s. - After the review, the one kit run of the registry, measured the same way four times: 5.91, 5.71, 5.70 and 5.82 s. The review added 64 single steps per exact GPU model to
FullSubmission, and two runs toSeedSensitivity's export. The registry used to run the kit twice, and now runs it once, socargo test -p henad-modelsspends about five seconds less. - Per model after the review, timed by a temporary test that was removed afterwards: sir 0.24 s, boids 0.43 s, game_of_life 0.30 s, ants 0.57 s, virus_network 0.06 s, team_assembly 0.03 s, gpu_game_of_life 0.13 s, gpu_sir 0.19 s, gpu_boids 0.51 s, gpu_ants 0.08 s. That is 2.5 s of checks, against 2.4 s before, and every CPU model reached 14 jobs.
- The rest of a process's time is the test binary's start and the device request.
[7.1] holds root overrides of the opt-level for henad-models and henad-compute in reserve before any shrinking of ThreadCount's size, and at these times neither is needed.
Proposed commits, in order:
test: grid thread-count test at 16 jobs,crates/henad-compute/src/cpu/grid_engine.rsalone, the named exception of [4.9.4].feat: testing kit, everything else.
Issues found & future directions¶
- The GPU checks build at the declared defaults. [4.8] has every built check shrink the sizes.
BaselineBuild,FullSubmission,SampledSliceand a GPU model'sActionsbuild at the defaults plus the overrides instead, as the old registry tests did. A watchdog trips on the time one command buffer runs, and a 128 by 128 grid never runs long enough to trip it, which would leaveFullSubmissionguarding nothing.DefaultsFitbounds the defaults to the baseline already. The other GPU builds shrink. - Additions to [4.8]'s API.
ModelCheck::DefaultSetup(the M6 review's check, whichDefaultsFitdoes not cover),ModelCheck::ALL,SkipReasonas an enum in place of a reason string,ModelReport::skip_reasonandthread_count_jobs(the job count [4.8] asks the report to record),SetReportascheck_model_set's return in place ofVec<ModelReport>,MIN_TICKS,Displayon the report types, andCheckSettings::thread_countspanicking on a low count not below the high one.CheckSettingsimplementsDefaultby hand, for 20 ticks and 1 and 7 threads.#[must_use]sits oncheck_modelandcheck_model_setat the review's request, an exception to AGENTS.md's rule against adding it by reflex. gpu_boidsnow skipsSeedSensitivity. [7]'s pin namesSameSeedandSamplingCadence, and the review adds the third. Two seeds of a model that does not replay exactly differ whether or not the seed reaches it.- Seeds and sizes of the built checks. The old tests built with the default seed at the declared sizes. The kit builds on seed 1 at the shrunk sizes, and
SeedSensitivitycompares seeds 1 and 2. - A GPU model on wasm32 skips every check that builds as
NativeOnly, where [4.8] names the three GPU checks.Simulationrefuses a GPU model in a browser, so the determinism checks could only fail there. assert_set_conformsprints to standard output. The workspace warns onprint_stdout, and the call carries anexpectwith its reason.- The grid search builds up to a dozen times. It knows no
rows_per_leaf, which is private to henad-compute, and readsparallel_jobsfrom builds. At 128 columns that is 12 small builds. - No self-test covers
Metadata,DefaultSetup,DefaultsFit,ParallelJobs,BaselineBuildorSampledSlice. Breaking one takes an entry built by hand, whichregister_*cannot produce, or a GPU model past the baseline. Correction after a later review:register_grid_modelaccepts a default outside its bounds, soDefaultSetupneeds no hand-built entry, anda_default_outside_its_bounds_fails_every_check_through_a_run_setupnow covers it. - A missing device under
HENAD_REQUIRE_GPUhas no self-test. Setting the variable inside a test changes it for every test of the process, andset_varisunsafeon edition 2024, which the workspace denies. check.sh's run sets it for every test, and a kit run without a device would fail there. authoring/testing.mdis not written. [8] gives it its own page with the template's test, and the template arrives in M10d. The determinism page carries the kit until then.- henad-models'
flumedev-dependency has no user, before and after M10a. cargo clippyfor wasm32 on henad-compute fails ondrop_non_dropingpu/sim_thread.rs, which M10a does not touch. check.sh runs no wasm32 clippy, and M10b's nightly atomics clippy will meet it.- Next. M10b, the facade, depends on M8, M9 and M10a, and re-exports the kit as
henad::testing.