Skip to content

Library M6: the programmatic API

M6 is the sixth of the eleven milestones of #48, which turns Henad into a library published on crates.io. A program now builds one model through a checked RunSetup, steps it as a Simulation, and reads its stats by label from a StatSample. run_spec runs a sweep or a search and plan_spec plans one, SweepOptions takes a spec file's execution table through apply_execution, and LoadedSpec reads a spec file whole. The CLI's benchmark moved into henad_explore::benchmark, streaming each repetition as it finishes, and the CLI's two exports step a Simulation. The performance gate caught run_sampled entering the pool once per sample, 7.6 times slower than M5's export at Game of Life 64², and it now enters the pool once per call. That gives its callback a Send bound the design did not have. A review of M6 before its commit then closed nine items, from a device-error test that passed without its error scopes to Build's refusals reaching the fault modal.

State before

48-library stood at abd6db1, M5's commit, with a clean tree. The CLI held its benchmark loops (bench_cpu, bench_gpu, run_gpu_rep), its export loops (stats_cpu, stats_gpu, export_final) and a private LoadedSpec in its binary. henad-explore exposed run_sweep and run_search with eight positional arguments each, a SweepOptions with Default, output_dir and dry_run, and a SweepRunOptions with a runtime field. A program built a model from positional values through ModelEntry::build, unchecked, and stepped the bare SimState.

What was done

The Simulation API (henad-compute)

henad_compute::simulation holds RunSetup, Simulation, SimulationViews, StatSample, SetupError and ExportError, as [4.2.2] gives them.

  • ModelEntry::setup() returns a setup at the defaults. set checks through check_value, set_text through parse_value, and act_at checks the id. from_parts checks positional values, a value count and each schedule entry, and from_replay refuses a run of another model.
  • RunSetup::build fires tick 0's actions before it returns, inside a pool entry on the CPU and inside catching_on on the GPU. A refused scheduled action becomes a Fault.
  • run_for and run_to test the schedule once per stretch and step to the next due tick in a bare loop. On the GPU they go through stepping::run_steps_acting under Fire::AfterStep.
  • Each CPU call enters rayon::scope once through cpu_call, with catching inside it on the worker. set_param stays unscoped, and write_state scopes only its prepare_view. A GPU call runs on the calling thread inside catching_on, and stats waits once after its blocking readback, so a fault outside the scopes reaches the call that raised it.
  • stepping::sample_stats records encode_stats_passes in place of the snapshot passes.
  • ParamValue gained From<f32>, From<u32> and From<bool>, and ModelEntry::schema() replaced henad-explore's model_schema.

The design has the GPU tick-0 actions fire "inside the build's catching_on scope". The factory owns that scope, so the actions fire in a second catching_on scope with BUILDING as its during, and a refusal or device error still comes back as the build's Fault.

The Simulation-shape gate

All five shapes ran against the M5 binary built from abd6db1 into its own target directory, release, rustc 1.97.1 (8bab26f4f 2026-07-14), on the 14-core M4 Pro (10 performance, 4 efficiency cores). A harness example in henad-cli, built in release against the M6 tree and deleted after the gates, ran shapes (ii) to (v). henad-models was unchanged between the two trees, so both sides step the same compiled kernels. Each configuration ran three rounds of the sequence (i), (ii), (iii), (iv), (iv), (iii), (ii), (i), five repetitions per run, seeds 1 to 5, and each shape pairs with the nearer (i) run, 30 pairs in all. No configuration came near the 1000 s cap. The longest, Game of Life at 64² with shape (iv), took 221 s. The measure is each repetition's elapsed time around run_for or the step loop, and elapsed_s from the M5 CLI's --json.

Configuration (ii) run_for (iii) step() in rayon::scope (iv) step() from main
Game of Life 64², 350,000 steps 0.999 1.009 12.21
Game of Life 256², 30,000 steps 0.975 0.975 2.05
Game of Life 1024², 10,000 steps 0.989 0.972 1.37
Game of Life 4096², 2,000 steps 1.003 1.015 1.07
SIR 64², 250,000 steps 1.000 1.008 8.71
gpu_game_of_life 1024², 20,000 steps, 1,000 global warm-up 0.981 against run_gpu_rep

Each cell is the median paired per-repetition ratio against shape (i). (ii) passes everywhere. (iii) and (iv) are measured for guide/library.md, which M11 writes. (iv) pays one pool entry per tick from outside the pool, about 15 µs, against a Game of Life tick of about 1.4 µs at 64². The 0.975 of (ii) at 256² is favourable and within the spread of the other rungs. It is flagged rather than read as a gain.

Shape (v), run_sampled(end, 1) writing a StatsWriter to a file, ran against a verbatim copy of the M5 CLI's stats_cpu loop in the same harness, ABBA, three rounds of five repetitions.

Configuration First build Rerun, cooled After the change
Game of Life 64², 50,000 steps 7.77 7.59 1.002
Game of Life 256², 5,000 steps 0.80 0.49
Game of Life 1024², 1,000 steps 0.88 0.73
Game of Life 4096², 100 steps 1.008 0.974
SIR 64², 50,000 steps 3.77 3.82 0.998

The first build entered the pool once per sample, as [4.2.2] says, and missed at both 64² rungs. A step of one rayon leaf runs inline on its caller, so M5's unscoped export loop never touched the pool there, and a pool entry per sample cost several ticks. The rerun after 90 s of cooling confirmed the miss. run_sampled on the CPU now runs the whole call inside one rayon::scope and calls on_sample inside it, outside the fault scopes. Its callback and its B take a WasmNotSend bound. The large wins at 256² and 1024² are not surprising: M5's export injected every parallel pass of every tick from outside the pool, and one scope removes those injects.

The benchmark and the exports

  • henad_explore::benchmark holds BenchmarkSettings, BenchmarkEvent, RepetitionReport, BenchmarkReport, BENCH_FIRE and run_benchmark. The timed shape is the CLI's: one rayon::scope per stretch holding a run_due and a step() per tick, Instant around it, the trailing run_due after the timer. catching, or catching_on on the GPU, wraps a whole repetition. Repetition i builds with seed.map(|seed| seed.wrapping_add(i)). A refused action is kept until the repetition ends and returned as a Fault, in place of the CLI's note.
  • The CLI prints the info line at Started, a whole rep line at each RepetitionFinished, and the population note on the CPU alone. actions.rs is gone, its test moved beside run_benchmark.
  • --export is run_to then write_state, and --export-stats is run_sampled with a callback that breaks on the first failed write.
  • params_by_id_json and scheduled_actions_json sit in henad_explore::output::details. The CLI's summary writes a choice as its index and the app's run details write it as its option name, as before, so the helper takes a ChoiceForm.

Sweeps, searches and the spec file

  • run_spec(model, gpu, spec, output, options, progress) dispatches on spec.search and on the output. plan_spec(model, gpu, spec, folder, options, progress) is the dry run. Both acquire a device for a GPU model handed none through sweep_device, which SweepRun now uses in place of BoundModel, and every path builds its ManifestRuntime from the context it steps on.
  • SweepInputs carries the folder and the dry-run flag in place of SweepOptions::output_dir and dry_run. run_sweep and run_search are gone rather than crate-private, since the four crate-private runners take an optional plan and cover every caller.
  • SweepOptions is #[non_exhaustive] without Default, with new(provenance), apply_execution, spec_source and provenance. SweepRunOptions likewise, with spec_source and no runtime. The private plan_spec of handle.rs is plan_for_start.
  • ExploreError::Device reports a device that cannot be acquired. ExploreError::NoOutput is gone, since nothing raises it.
  • LoadedSpec (read, parse, to_toml) replaced the CLI's private type, and SpecFileError::Write wraps the TOML writer's error. The app's Save spec writes through to_toml.
  • ResultSet::schema_matches, plan and replay take a ModelSchema.
  • The thirteen free functions of [7]'s M6 entry are pub(crate). None had a caller outside henad-explore.
  • install_panic_hook left run_into_directory, run_in_memory, their two search twins, the removed run_sweep and run_search, and PumpedSweep. support::sweep and sweep_with install it.
  • The CLI applies a spec's table, then each of --concurrent, --memory and --gpu-memory it was given. The app's Build checks its fields through RunSetup::from_parts, and the review below moved that check into Build's enabled state.

Equivalence and the instrument gate

  • --export-stats at seed 1, 40 steps, --stats-every 1 and 3, wrote identical bytes from the M5 and M6 binaries for all ten models, without actions and with the first declared action at ticks 0 and 7. gpu_boids compared its tick column.
  • --export at seed 1, 3 warm-up and 20 steps wrote identical bytes for the six CPU models, SIR with outbreaks at 0, 7 and 23 and Game of Life with randomise@0, clear@7 and randomise@9.
  • henad-cli sir --export out.csv --steps 0 --act seed_outbreak@0 wrote the bytes the v0.2.0 binary writes, which differ from the export without the action.
  • The --json lines of six benchmarks matched the M5 binary's apart from the timings and engine_version: SIR with and without --seed, with actions at 0, 60 and past the end, Virus on a Network on its geometric network, boids, gpu_sir with actions and gpu_boids without --seed. The stderr lines matched up to their timings.
  • A Game of Life benchmark killed with SIGKILL after 3 s had printed the info line and 11 rep lines, the 11 repetitions its stderr reported finished. A gpu_game_of_life one printed 5 of 5.

The instrument gate ran the M5 binary's benchmark against M6's over G, --json --seed 1, each configuration in rounds of M5, M6, M6, M5, each M6 run paired repetition by repetition with the M5 run beside it. The CPU rungs ran bench_matrix's 1000 steps with a 200-step warm-up, and the GPU rungs its 100,000 steps with a 20,000-step warm-up and a 10,000-step global warm-up. Three GPU rungs ran fewer steps to fit the cap, as named in the table, since gpu_boids at 1M agents and 100,000 steps did not finish one repetition in 9 minutes. The longest configuration, boids at 50,000 agents, took 590 s.

Configuration Repetitions × runs M5 median M6 median Median ratio
game_of_life 64² 5 × 6 1.64 ms 1.67 ms 1.014
game_of_life 1024² 5 × 6 50.9 ms 50.0 ms 0.979
game_of_life 4096² 3 × 6 227 ms 219 ms 0.946
sir 64² 5 × 6 1.95 ms 2.04 ms 1.006
sir 1024² 5 × 6 84.1 ms 80.3 ms 0.995
boids 1,000 3 × 6 138 ms 122 ms 0.858, rerun 0.863
boids 50,000 3 × 4 20.9 s 20.6 s 1.009
ants 10,000 3 × 6 454 ms 456 ms 1.006
ants 1,000,000 3 × 4 6.24 s 6.21 s 1.005
virus_network 10,000 5 × 6 30.3 ms 32.5 ms 1.072, rerun 1.036
team_assembly, defaults 5 × 6 0.52 ms 0.52 ms 0.982
gpu_game_of_life 256² 3 × 4 2.80 s 2.82 s 1.007
gpu_game_of_life 4096² 3 × 4 3.92 s 3.85 s 0.994
gpu_sir 1024² 3 × 4 11.5 s 11.4 s 1.004
gpu_boids 10,000, 20,000 steps 3 × 4 3.54 s 3.55 s 1.000
gpu_boids 100,000, 200 steps 3 × 4 2.80 s 2.79 s 1.000
gpu_ants 100,000, 2,000 steps 3 × 4 134 ms 133 ms 0.991

virus_network missed at 1.072 and passed at 1.036 on a rerun after 90 s of cooling. Its repetitions last about 30 ms, and the M5 median moved from 30 ms to 41 ms between the two runs, so the machine's noise there is larger than the margin. Both runs sit above 1, and a cost of a few percent on that rung cannot be ruled out from these numbers. boids at 1,000 agents ran 14% faster on both runs. The timed loop is the same code moved into henad-explore, and the gain is unexplained. It is flagged as surprising, not read as a finding.

scripts/bench_matrix.py --models game_of_life --reps 2 ran once against each binary, 28 rows each. Every column but the timings matched, including the four rows at grid size 0 that both binaries refuse with the same error. scripts/bench_matrix.py --dry-run printed the same matrix through both.

Tests

New: a_simulation_follows_the_run_cursor (replay.rs), a_tick_zero_action_is_in_the_state_build_returns, a_setup_refuses_a_value_of_the_wrong_kind, a_setup_refuses_a_wrong_value_count, an_unsuffixed_literal_sets_a_parameter, views_are_prepared_at_the_read and a_gpu_simulation_reports_its_own_device_error (henad-models' tests/simulation.rs), plan_spec_matches_a_dry_run (sweep.rs), a_dry_run_counts_the_runs_a_search_resume_skips (tests/search.rs), an_explicit_concurrent_auto_overrides_the_spec_table (henad-cli's explore.rs) and a_benchmark_fires_each_action_before_its_step (benchmark.rs).

  • a_gpu_simulation_reports_its_own_device_error wraps gpu_game_of_life's state so that a live parameter edit asks the device for a buffer over its limit. The review below rewrote it. Its first form raised the error from run_for, whose wait reports the sink, and it passed without the error scopes.
  • a_benchmark_fires_each_action_before_its_step registers a counting grid model whose action records the cell count it sees, and checks [0, 2, 3, 5] per repetition for actions at 0, 2, 3, 5 and 9 with 2 warm-up and 3 timed steps. It also checks the seeds base + i, the event order, and None seeds for a setup on the default seed.
  • Moved: a_gpu_benchmark_rep_fires_every_action_once is a_gpu_benchmark_repetition_fires_every_action_once beside run_benchmark, and compares a repetition with a Simulation stepped over the same ticks. a_timed_run_fires_the_ticks_the_cpu_loop_times moved there from actions.rs.
  • Changed: sampling_cadence_does_not_change_the_trajectory read back a display that stepping::sample_stats drew as a side effect. A sample records the stats passes alone now, and the test records the snapshot passes once before it reads the view.
  • Rewritten onto Simulation: the five CLI tests that called stats_cpu, stats_gpu and run_gpu_rep, through a write_series helper the export uses too.
  • Rewritten onto run_spec and plan_spec: every henad-explore test that called run_sweep or run_search. a_dry_run_counts_the_runs_a_resume_skips_and_changes_nothing keeps its counts of 3 and 3 and its CommitChanged, which M7 replaces.
  • Dropped: the half of a_resume_refuses_another_search_or_a_sweep that ran a search spec as a sweep, since no public path does that any more, and the rep_seed assertion of existing_invocations_keep_their_mode, which the benchmark test covers.

The M6 review

The maintainer reviewed M6 before its commit and listed nine items.

  1. The device-error test passed without its scopes. run_for ends in stepping::wait, which reports whatever the sink holds, so a validation error reached it with or without catching_on. The test now raises the oversized buffer from the wrapped state's set_param, a call that submits nothing and waits for nothing, and expects SetupError::Fault holding FaultKind::Device. A healthy simulation on the same context then steps, and the sink stays empty. With catching_on replaced by catching in Simulation::set_param, the test failed with ReloadOnly and the error in the sink. Simulation::set_param used to refuse a reload-only parameter from its descriptor before the call reached the state, and every GPU parameter is reload-only. It now leaves that refusal to the state, which the CPU ParamStore and both GPU engines already make.
  2. Build's refusals. AppState::build_setup runs RunSetup::from_parts over the fields, Playback's build_refusal shows a SetupError as Build's disabled reason through setup_message, after the shortfalls and INVALID_SEED, and reset_simulation builds nothing while the check fails. open_run refuses a replay RunSetup::from_replay refuses, with the reason, before it touches the fields. every_default_setup_passes_the_checks_of_from_parts runs from_parts over each example model's defaults. The sliders clamp every edit, so the Parameters panel cannot hold a refused value. A temporary patch, reverted before the hand-over, started SIR with an infection rate of 2. Driven through the egui MCP server, Build read disabled with the tooltip "Parameter 'infection_rate': 2 is outside 0..=1", and dragging the slider into range re-enabled it and built the model. The app's saved state was backed up before and restored after.
  3. CHANGELOG. run_sweep and run_search are under Removed. ExploreError::Device, the panic hook and sample_stats dropping the display pass are marked "Breaking:". SweepRunOptions::runtime has its migration through GpuContext::with_runtime_info. A line says the CLI exits 1 with a message on a model panic, where it exited 101, and treats a refused scheduled action as an error. --export of a model with no grid or point view fails as it did in 0.2.0, from the entry's topology_hint before anything is built or written.
  4. CPU sweeps record no adapter. sweep_device returns None for an entry without gpu_needs(), whatever context it is handed, so a CPU sweep from the CLI, the app or a program records no adapter. henad-cli's CPU sweeps recorded the machine's adapter before, and the CHANGELOG says so. a_sweep_records_the_adapter_of_the_context_it_steps_on checks that a GPU sweep handed a context records its adapter, and that a CPU sweep handed the same context records none.
  5. Tests. a_panicking_sample_callback_unwinds_as_a_panic catches the callback's own panic message out of run_sampled. a_break_at_the_first_sample_leaves_the_tick breaks at tick 3 and finds the simulation still there. A const in simulation.rs asserts Simulation: Send on native.
  6. Schedule::fire_ticks builds its windows with saturating_add.
  7. ExecutionBudget, MAX_AUTO_GPU_TRACKS and LARGE_GPU_POPULATION are pub(crate). SummaryError stays public, since the public OutputError::Summary carries it and a crate-private type there trips private_interfaces. SearchPlanError::NotASweep is gone, and SweepPreparation::new debug-asserts that its spec has no [search] table.
  8. Docs. The simulation module and Simulation name the calls that enter the pool, and set_param among those that do not. run_sampled says its callback holds a pool worker. The benchmark's module doc says the warm-up and the timed steps each have a scope, and the history in it and in run_benchmark is gone. BenchmarkSettings::repetitions, both populations of RepetitionReport, BenchmarkReport's fields and RepetitionFinished have docs, and BenchmarkEvent, RepetitionReport and BenchmarkReport are #[non_exhaustive], with a wildcard arm in the CLI. LoadedSpec::parse says it refuses a table file, and that an inline table_text is read. AGENTS.md drops "as bench_cpu does", describes build_setup and Save spec through LoadedSpec::to_toml, and rewraps the panic-hook passage. SweepRunOptions' doc is rewrapped.
  9. The design. [4.2.2]'s run_sampled signature and its bullet in dev-docs/48-library/design-v6.md give the shipped rule, one pool entry per call with the callback inside it and Send, and cite this record's gate. dev-docs/ is gitignored, so that edit stays out of the commit.

Docs

reference/cli.md says that both exports step a Simulation, that tick 0's actions fire before the first step, and that a CPU export enters the pool once. AGENTS.md describes simulation.rs, run_spec and plan_spec, LoadedSpec, the panic-hook rule, benchmark.rs and output/details.rs, and drops run_sweep, run_search, BoundModel and actions.rs.

Edited tree

.
├── AGENTS.md                                 ~ simulation.rs, run_spec and plan_spec, the benchmark, the exports
├── CHANGELOG.md                              ~ M6's Added, Changed and Removed
├── zensical.toml                             ~ nav entry #35
├── crates/
│   ├── henad-core/src/
│   │   ├── params.rs                         ~ From<f32>, From<u32>, From<bool> for ParamValue
│   │   └── action.rs                         ~ fire_ticks saturates
│   ├── henad-compute/src/
│   │   ├── lib.rs                            ~ pub mod simulation
│   │   ├── simulation.rs                     + RunSetup, Simulation, StatSample, SetupError, ExportError
│   │   ├── entry/mod.rs                      ~ setup, schema
│   │   └── gpu/stepping.rs                   ~ sample_stats over encode_stats_passes
│   ├── henad-models/src/tests/
│   │   ├── mod.rs                            ~ mod simulation
│   │   ├── registry.rs                       ~ every_default_setup_passes_the_checks_of_from_parts
│   │   └── simulation.rs                     + the setup and simulation tests
│   ├── henad-explore/src/
│   │   ├── lib.rs                            ~ pub mod benchmark
│   │   ├── benchmark.rs                      + run_benchmark and its tests
│   │   ├── sweep.rs                          ~ run_spec, plan_spec, sweep_device, SweepOptions::new, apply_execution
│   │   ├── search_run.rs                     ~ run_search and NotASweep gone, the folder from the inputs
│   │   ├── handle.rs                         ~ SweepRunOptions::new, spec_source, plan_for_start, BoundModel gone
│   │   ├── pumped.rs                         ~ options through new, no panic hook
│   │   ├── spec_file.rs                      ~ LoadedSpec, SpecFileError::Write
│   │   ├── result_set.rs                     ~ ModelSchema in place of &ModelEntry
│   │   ├── schema.rs                         ~ model_schema gone
│   │   ├── probe.rs, exec/mod.rs             ~ pub(crate)
│   │   ├── output/details.rs                 + params_by_id_json, scheduled_actions_json
│   │   ├── output/{manifest,memory,mod,read,runs_csv,series_csv,summary_csv,search_tables}.rs   ~ pub(crate)
│   │   └── tests/                            ~ support.rs, replay.rs, search.rs, gpu.rs, resume.rs and the other call sites
│   ├── henad-cli/src/
│   │   ├── main.rs                           ~ benchmark through run_benchmark, exports through Simulation
│   │   ├── actions.rs                        - moved into henad-explore's benchmark.rs
│   │   ├── explore.rs                        ~ LoadedSpec, run_spec and plan_spec, flags after the table
│   │   └── json_report.rs                    ~ the shared JSON helpers
│   └── henad-app/src/
│       ├── state.rs                          ~ build_setup, setup_message, open_run through from_replay
│       ├── ui/playback.rs                    ~ a SetupError as Build's disabled reason
│       └── ui/{sweep,results,export}/        ~ SweepRunOptions::new, LoadedSpec::to_toml, entry.schema(), details helpers
└── docs/
    ├── reference/cli.md                      ~ the exports through Simulation
    └── developing/agent-record/20261001-35-library-api.md   +

State after

M6 is implemented and uncommitted on 48-library, on top of abd6db1. The workspace version stays 0.2.0.

  • HENAD_REQUIRE_GPU=1 ./check.sh passes: 1006 tests, none failed, with the wasm32 typecheck, the packaging, cargo-deny and docs steps and the web build. M5 passed at 991, and M6 before its review at 1002. The first run before the review stopped at rustdoc, on public docs linking to the now crate-private choose_layout and on a link from SweepRunOptions to SweepOptions outside its scope.
  • The review's changes touch no measured path. run_benchmark, run_for and run_sampled are as the gates timed them, and the gates were not rerun.
  • uv run --locked zensical build passes with its path checks.
  • Every M6 gate of [7.1] passes against the M5 binary, with run_sampled changed after its miss and virus_network passing on its rerun.
  • --export-stats and --export write the M5 binary's bytes, and the 0.2.0 outbreak export its bytes.

Proposed commit message: feat: programmatic API

Issues found & future directions

  • run_sampled's callback is Send. [4.2.2] wanted it called outside the pool with no Send bound, at one pool entry per sample. That shape missed the gate by 7.6 times at Game of Life 64², and the bound is the cost of one pool entry per call. A host whose callback holds something that is not Send steps through run_to and stats in its own loop. The design's [4.2.2] now says so, and guide/library.md takes the same wording when M11 writes it.
  • The app cannot show a refused value on its own. Every Parameters widget clamps, so Build's new reason appears only for fields set from elsewhere. Build checks them all the same, and a future loader that sets the fields directly meets the check.
  • A loop of step() from main costs a pool entry per tick. Shape (iv) is 12 times shape (i) at Game of Life 64² and 1.07 times at 4096². Simulation::step's docs say to wrap such a loop in rayon::scope or pool.install, and shape (iii) shows that this recovers shape (i).
  • The GPU tick-0 actions fire in a scope of their own after the factory's, with BUILDING as their during. Moving them into the factory's scope would need the schedule inside the factory.
  • Reduced rungs. bench_matrix's 100,000 GPU steps would hold gpu_boids at 1M agents for over 9 minutes a repetition, past the cap. The GPU rungs that ran with fewer steps are named in the instrument table.
  • Next. M7 (provenance) depends on M3 and M5 and can start. M8 and M9 wait for M7.

Manual notes (human)