Library M6: the programmatic API¶
M6 is the sixth of the eleven milestones of #48, which turns Henad into a library published on crates.io. A program now builds one model through a checked
RunSetup, steps it as aSimulation, and reads its stats by label from aStatSample.run_specruns a sweep or a search andplan_specplans one,SweepOptionstakes a spec file's execution table throughapply_execution, andLoadedSpecreads a spec file whole. The CLI's benchmark moved intohenad_explore::benchmark, streaming each repetition as it finishes, and the CLI's two exports step aSimulation. The performance gate caughtrun_sampledentering the pool once per sample, 7.6 times slower than M5's export at Game of Life 64², and it now enters the pool once per call. That gives its callback aSendbound the design did not have. A review of M6 before its commit then closed nine items, from a device-error test that passed without its error scopes to Build's refusals reaching the fault modal.
State before¶
48-library stood at abd6db1, M5's commit, with a clean tree.
The CLI held its benchmark loops (bench_cpu, bench_gpu, run_gpu_rep), its export loops (stats_cpu, stats_gpu, export_final) and a private LoadedSpec in its binary.
henad-explore exposed run_sweep and run_search with eight positional arguments each, a SweepOptions with Default, output_dir and dry_run, and a SweepRunOptions with a runtime field.
A program built a model from positional values through ModelEntry::build, unchecked, and stepped the bare SimState.
What was done¶
The Simulation API (henad-compute)¶
henad_compute::simulation holds RunSetup, Simulation, SimulationViews, StatSample, SetupError and ExportError, as [4.2.2] gives them.
ModelEntry::setup()returns a setup at the defaults.setchecks throughcheck_value,set_textthroughparse_value, andact_atchecks the id.from_partschecks positional values, a value count and each schedule entry, andfrom_replayrefuses a run of another model.RunSetup::buildfires tick 0's actions before it returns, inside a pool entry on the CPU and insidecatching_onon the GPU. A refused scheduled action becomes aFault.run_forandrun_totest the schedule once per stretch and step to the next due tick in a bare loop. On the GPU they go throughstepping::run_steps_actingunderFire::AfterStep.- Each CPU call enters
rayon::scopeonce throughcpu_call, withcatchinginside it on the worker.set_paramstays unscoped, andwrite_statescopes only itsprepare_view. A GPU call runs on the calling thread insidecatching_on, andstatswaits once after its blocking readback, so a fault outside the scopes reaches the call that raised it. stepping::sample_statsrecordsencode_stats_passesin place of the snapshot passes.ParamValuegainedFrom<f32>,From<u32>andFrom<bool>, andModelEntry::schema()replaced henad-explore'smodel_schema.
The design has the GPU tick-0 actions fire "inside the build's catching_on scope". The factory owns that scope, so the actions fire in a second catching_on scope with BUILDING as its during, and a refusal or device error still comes back as the build's Fault.
The Simulation-shape gate¶
All five shapes ran against the M5 binary built from abd6db1 into its own target directory, release, rustc 1.97.1 (8bab26f4f 2026-07-14), on the 14-core M4 Pro (10 performance, 4 efficiency cores).
A harness example in henad-cli, built in release against the M6 tree and deleted after the gates, ran shapes (ii) to (v).
henad-models was unchanged between the two trees, so both sides step the same compiled kernels.
Each configuration ran three rounds of the sequence (i), (ii), (iii), (iv), (iv), (iii), (ii), (i), five repetitions per run, seeds 1 to 5, and each shape pairs with the nearer (i) run, 30 pairs in all.
No configuration came near the 1000 s cap. The longest, Game of Life at 64² with shape (iv), took 221 s.
The measure is each repetition's elapsed time around run_for or the step loop, and elapsed_s from the M5 CLI's --json.
| Configuration | (ii) run_for |
(iii) step() in rayon::scope |
(iv) step() from main |
|---|---|---|---|
| Game of Life 64², 350,000 steps | 0.999 | 1.009 | 12.21 |
| Game of Life 256², 30,000 steps | 0.975 | 0.975 | 2.05 |
| Game of Life 1024², 10,000 steps | 0.989 | 0.972 | 1.37 |
| Game of Life 4096², 2,000 steps | 1.003 | 1.015 | 1.07 |
| SIR 64², 250,000 steps | 1.000 | 1.008 | 8.71 |
| gpu_game_of_life 1024², 20,000 steps, 1,000 global warm-up | 0.981 against run_gpu_rep |
Each cell is the median paired per-repetition ratio against shape (i). (ii) passes everywhere.
(iii) and (iv) are measured for guide/library.md, which M11 writes.
(iv) pays one pool entry per tick from outside the pool, about 15 µs, against a Game of Life tick of about 1.4 µs at 64².
The 0.975 of (ii) at 256² is favourable and within the spread of the other rungs. It is flagged rather than read as a gain.
Shape (v), run_sampled(end, 1) writing a StatsWriter to a file, ran against a verbatim copy of the M5 CLI's stats_cpu loop in the same harness, ABBA, three rounds of five repetitions.
| Configuration | First build | Rerun, cooled | After the change |
|---|---|---|---|
| Game of Life 64², 50,000 steps | 7.77 | 7.59 | 1.002 |
| Game of Life 256², 5,000 steps | 0.80 | 0.49 | |
| Game of Life 1024², 1,000 steps | 0.88 | 0.73 | |
| Game of Life 4096², 100 steps | 1.008 | 0.974 | |
| SIR 64², 50,000 steps | 3.77 | 3.82 | 0.998 |
The first build entered the pool once per sample, as [4.2.2] says, and missed at both 64² rungs.
A step of one rayon leaf runs inline on its caller, so M5's unscoped export loop never touched the pool there, and a pool entry per sample cost several ticks.
The rerun after 90 s of cooling confirmed the miss.
run_sampled on the CPU now runs the whole call inside one rayon::scope and calls on_sample inside it, outside the fault scopes. Its callback and its B take a WasmNotSend bound.
The large wins at 256² and 1024² are not surprising: M5's export injected every parallel pass of every tick from outside the pool, and one scope removes those injects.
The benchmark and the exports¶
henad_explore::benchmarkholdsBenchmarkSettings,BenchmarkEvent,RepetitionReport,BenchmarkReport,BENCH_FIREandrun_benchmark. The timed shape is the CLI's: onerayon::scopeper stretch holding arun_dueand astep()per tick,Instantaround it, the trailingrun_dueafter the timer.catching, orcatching_onon the GPU, wraps a whole repetition. Repetitionibuilds withseed.map(|seed| seed.wrapping_add(i)). A refused action is kept until the repetition ends and returned as aFault, in place of the CLI's note.- The CLI prints the
infoline atStarted, a wholerepline at eachRepetitionFinished, and the population note on the CPU alone.actions.rsis gone, its test moved besiderun_benchmark. --exportisrun_tothenwrite_state, and--export-statsisrun_sampledwith a callback that breaks on the first failed write.params_by_id_jsonandscheduled_actions_jsonsit inhenad_explore::output::details. The CLI's summary writes a choice as its index and the app's run details write it as its option name, as before, so the helper takes aChoiceForm.
Sweeps, searches and the spec file¶
run_spec(model, gpu, spec, output, options, progress)dispatches onspec.searchand on the output.plan_spec(model, gpu, spec, folder, options, progress)is the dry run. Both acquire a device for a GPU model handed none throughsweep_device, whichSweepRunnow uses in place ofBoundModel, and every path builds itsManifestRuntimefrom the context it steps on.SweepInputscarries the folder and the dry-run flag in place ofSweepOptions::output_diranddry_run.run_sweepandrun_searchare gone rather than crate-private, since the four crate-private runners take an optional plan and cover every caller.SweepOptionsis#[non_exhaustive]withoutDefault, withnew(provenance),apply_execution,spec_sourceandprovenance.SweepRunOptionslikewise, withspec_sourceand noruntime. The privateplan_specofhandle.rsisplan_for_start.ExploreError::Devicereports a device that cannot be acquired.ExploreError::NoOutputis gone, since nothing raises it.LoadedSpec(read,parse,to_toml) replaced the CLI's private type, andSpecFileError::Writewraps the TOML writer's error. The app's Save spec writes throughto_toml.ResultSet::schema_matches,planandreplaytake aModelSchema.- The thirteen free functions of [7]'s M6 entry are
pub(crate). None had a caller outside henad-explore. install_panic_hookleftrun_into_directory,run_in_memory, their two search twins, the removedrun_sweepandrun_search, andPumpedSweep.support::sweepandsweep_withinstall it.- The CLI applies a spec's table, then each of
--concurrent,--memoryand--gpu-memoryit was given. The app's Build checks its fields throughRunSetup::from_parts, and the review below moved that check into Build's enabled state.
Equivalence and the instrument gate¶
--export-statsat seed 1, 40 steps,--stats-every1 and 3, wrote identical bytes from the M5 and M6 binaries for all ten models, without actions and with the first declared action at ticks 0 and 7. gpu_boids compared its tick column.--exportat seed 1, 3 warm-up and 20 steps wrote identical bytes for the six CPU models, SIR with outbreaks at 0, 7 and 23 and Game of Life withrandomise@0,clear@7andrandomise@9.henad-cli sir --export out.csv --steps 0 --act seed_outbreak@0wrote the bytes the v0.2.0 binary writes, which differ from the export without the action.- The
--jsonlines of six benchmarks matched the M5 binary's apart from the timings andengine_version: SIR with and without--seed, with actions at 0, 60 and past the end, Virus on a Network on its geometric network, boids, gpu_sir with actions and gpu_boids without--seed. The stderr lines matched up to their timings. - A Game of Life benchmark killed with SIGKILL after 3 s had printed the
infoline and 11replines, the 11 repetitions its stderr reported finished. A gpu_game_of_life one printed 5 of 5.
The instrument gate ran the M5 binary's benchmark against M6's over G, --json --seed 1, each configuration in rounds of M5, M6, M6, M5, each M6 run paired repetition by repetition with the M5 run beside it.
The CPU rungs ran bench_matrix's 1000 steps with a 200-step warm-up, and the GPU rungs its 100,000 steps with a 20,000-step warm-up and a 10,000-step global warm-up.
Three GPU rungs ran fewer steps to fit the cap, as named in the table, since gpu_boids at 1M agents and 100,000 steps did not finish one repetition in 9 minutes.
The longest configuration, boids at 50,000 agents, took 590 s.
| Configuration | Repetitions × runs | M5 median | M6 median | Median ratio |
|---|---|---|---|---|
| game_of_life 64² | 5 × 6 | 1.64 ms | 1.67 ms | 1.014 |
| game_of_life 1024² | 5 × 6 | 50.9 ms | 50.0 ms | 0.979 |
| game_of_life 4096² | 3 × 6 | 227 ms | 219 ms | 0.946 |
| sir 64² | 5 × 6 | 1.95 ms | 2.04 ms | 1.006 |
| sir 1024² | 5 × 6 | 84.1 ms | 80.3 ms | 0.995 |
| boids 1,000 | 3 × 6 | 138 ms | 122 ms | 0.858, rerun 0.863 |
| boids 50,000 | 3 × 4 | 20.9 s | 20.6 s | 1.009 |
| ants 10,000 | 3 × 6 | 454 ms | 456 ms | 1.006 |
| ants 1,000,000 | 3 × 4 | 6.24 s | 6.21 s | 1.005 |
| virus_network 10,000 | 5 × 6 | 30.3 ms | 32.5 ms | 1.072, rerun 1.036 |
| team_assembly, defaults | 5 × 6 | 0.52 ms | 0.52 ms | 0.982 |
| gpu_game_of_life 256² | 3 × 4 | 2.80 s | 2.82 s | 1.007 |
| gpu_game_of_life 4096² | 3 × 4 | 3.92 s | 3.85 s | 0.994 |
| gpu_sir 1024² | 3 × 4 | 11.5 s | 11.4 s | 1.004 |
| gpu_boids 10,000, 20,000 steps | 3 × 4 | 3.54 s | 3.55 s | 1.000 |
| gpu_boids 100,000, 200 steps | 3 × 4 | 2.80 s | 2.79 s | 1.000 |
| gpu_ants 100,000, 2,000 steps | 3 × 4 | 134 ms | 133 ms | 0.991 |
virus_network missed at 1.072 and passed at 1.036 on a rerun after 90 s of cooling. Its repetitions last about 30 ms, and the M5 median moved from 30 ms to 41 ms between the two runs, so the machine's noise there is larger than the margin. Both runs sit above 1, and a cost of a few percent on that rung cannot be ruled out from these numbers. boids at 1,000 agents ran 14% faster on both runs. The timed loop is the same code moved into henad-explore, and the gain is unexplained. It is flagged as surprising, not read as a finding.
scripts/bench_matrix.py --models game_of_life --reps 2 ran once against each binary, 28 rows each.
Every column but the timings matched, including the four rows at grid size 0 that both binaries refuse with the same error.
scripts/bench_matrix.py --dry-run printed the same matrix through both.
Tests¶
New: a_simulation_follows_the_run_cursor (replay.rs), a_tick_zero_action_is_in_the_state_build_returns, a_setup_refuses_a_value_of_the_wrong_kind, a_setup_refuses_a_wrong_value_count, an_unsuffixed_literal_sets_a_parameter, views_are_prepared_at_the_read and a_gpu_simulation_reports_its_own_device_error (henad-models' tests/simulation.rs), plan_spec_matches_a_dry_run (sweep.rs), a_dry_run_counts_the_runs_a_search_resume_skips (tests/search.rs), an_explicit_concurrent_auto_overrides_the_spec_table (henad-cli's explore.rs) and a_benchmark_fires_each_action_before_its_step (benchmark.rs).
a_gpu_simulation_reports_its_own_device_errorwraps gpu_game_of_life's state so that a live parameter edit asks the device for a buffer over its limit. The review below rewrote it. Its first form raised the error fromrun_for, whose wait reports the sink, and it passed without the error scopes.a_benchmark_fires_each_action_before_its_stepregisters a counting grid model whose action records the cell count it sees, and checks[0, 2, 3, 5]per repetition for actions at 0, 2, 3, 5 and 9 with 2 warm-up and 3 timed steps. It also checks the seedsbase + i, the event order, andNoneseeds for a setup on the default seed.- Moved:
a_gpu_benchmark_rep_fires_every_action_onceisa_gpu_benchmark_repetition_fires_every_action_oncebesiderun_benchmark, and compares a repetition with aSimulationstepped over the same ticks.a_timed_run_fires_the_ticks_the_cpu_loop_timesmoved there fromactions.rs. - Changed:
sampling_cadence_does_not_change_the_trajectoryread back a display thatstepping::sample_statsdrew as a side effect. A sample records the stats passes alone now, and the test records the snapshot passes once before it reads the view. - Rewritten onto
Simulation: the five CLI tests that calledstats_cpu,stats_gpuandrun_gpu_rep, through awrite_serieshelper the export uses too. - Rewritten onto
run_specandplan_spec: every henad-explore test that calledrun_sweeporrun_search.a_dry_run_counts_the_runs_a_resume_skips_and_changes_nothingkeeps its counts of 3 and 3 and itsCommitChanged, which M7 replaces. - Dropped: the half of
a_resume_refuses_another_search_or_a_sweepthat ran a search spec as a sweep, since no public path does that any more, and therep_seedassertion ofexisting_invocations_keep_their_mode, which the benchmark test covers.
The M6 review¶
The maintainer reviewed M6 before its commit and listed nine items.
- The device-error test passed without its scopes.
run_forends instepping::wait, which reports whatever the sink holds, so a validation error reached it with or withoutcatching_on. The test now raises the oversized buffer from the wrapped state'sset_param, a call that submits nothing and waits for nothing, and expectsSetupError::FaultholdingFaultKind::Device. A healthy simulation on the same context then steps, and the sink stays empty. Withcatching_onreplaced bycatchinginSimulation::set_param, the test failed withReloadOnlyand the error in the sink.Simulation::set_paramused to refuse a reload-only parameter from its descriptor before the call reached the state, and every GPU parameter is reload-only. It now leaves that refusal to the state, which the CPUParamStoreand both GPU engines already make. - Build's refusals.
AppState::build_setuprunsRunSetup::from_partsover the fields, Playback'sbuild_refusalshows aSetupErroras Build's disabled reason throughsetup_message, after the shortfalls andINVALID_SEED, andreset_simulationbuilds nothing while the check fails.open_runrefuses a replayRunSetup::from_replayrefuses, with the reason, before it touches the fields.every_default_setup_passes_the_checks_of_from_partsrunsfrom_partsover each example model's defaults. The sliders clamp every edit, so the Parameters panel cannot hold a refused value. A temporary patch, reverted before the hand-over, started SIR with an infection rate of 2. Driven through the egui MCP server, Build read disabled with the tooltip "Parameter 'infection_rate': 2 is outside 0..=1", and dragging the slider into range re-enabled it and built the model. The app's saved state was backed up before and restored after. - CHANGELOG.
run_sweepandrun_searchare under Removed.ExploreError::Device, the panic hook andsample_statsdropping the display pass are marked "Breaking:".SweepRunOptions::runtimehas its migration throughGpuContext::with_runtime_info. A line says the CLI exits 1 with a message on a model panic, where it exited 101, and treats a refused scheduled action as an error.--exportof a model with no grid or point view fails as it did in 0.2.0, from the entry'stopology_hintbefore anything is built or written. - CPU sweeps record no adapter.
sweep_devicereturnsNonefor an entry withoutgpu_needs(), whatever context it is handed, so a CPU sweep from the CLI, the app or a program records no adapter.henad-cli's CPU sweeps recorded the machine's adapter before, and the CHANGELOG says so.a_sweep_records_the_adapter_of_the_context_it_steps_onchecks that a GPU sweep handed a context records its adapter, and that a CPU sweep handed the same context records none. - Tests.
a_panicking_sample_callback_unwinds_as_a_paniccatches the callback's own panic message out ofrun_sampled.a_break_at_the_first_sample_leaves_the_tickbreaks at tick 3 and finds the simulation still there. Aconstinsimulation.rsassertsSimulation: Sendon native. Schedule::fire_ticksbuilds its windows withsaturating_add.ExecutionBudget,MAX_AUTO_GPU_TRACKSandLARGE_GPU_POPULATIONarepub(crate).SummaryErrorstays public, since the publicOutputError::Summarycarries it and a crate-private type there tripsprivate_interfaces.SearchPlanError::NotASweepis gone, andSweepPreparation::newdebug-asserts that its spec has no[search]table.- Docs. The
simulationmodule andSimulationname the calls that enter the pool, andset_paramamong those that do not.run_sampledsays its callback holds a pool worker. The benchmark's module doc says the warm-up and the timed steps each have a scope, and the history in it and inrun_benchmarkis gone.BenchmarkSettings::repetitions, both populations ofRepetitionReport,BenchmarkReport's fields andRepetitionFinishedhave docs, andBenchmarkEvent,RepetitionReportandBenchmarkReportare#[non_exhaustive], with a wildcard arm in the CLI.LoadedSpec::parsesays it refuses a table file, and that an inlinetable_textis read. AGENTS.md drops "asbench_cpudoes", describesbuild_setupand Save spec throughLoadedSpec::to_toml, and rewraps the panic-hook passage.SweepRunOptions' doc is rewrapped. - The design. [4.2.2]'s
run_sampledsignature and its bullet in dev-docs/48-library/design-v6.md give the shipped rule, one pool entry per call with the callback inside it andSend, and cite this record's gate. dev-docs/ is gitignored, so that edit stays out of the commit.
Docs¶
reference/cli.md says that both exports step a Simulation, that tick 0's actions fire before the first step, and that a CPU export enters the pool once.
AGENTS.md describes simulation.rs, run_spec and plan_spec, LoadedSpec, the panic-hook rule, benchmark.rs and output/details.rs, and drops run_sweep, run_search, BoundModel and actions.rs.
Edited tree¶
.
├── AGENTS.md ~ simulation.rs, run_spec and plan_spec, the benchmark, the exports
├── CHANGELOG.md ~ M6's Added, Changed and Removed
├── zensical.toml ~ nav entry #35
├── crates/
│ ├── henad-core/src/
│ │ ├── params.rs ~ From<f32>, From<u32>, From<bool> for ParamValue
│ │ └── action.rs ~ fire_ticks saturates
│ ├── henad-compute/src/
│ │ ├── lib.rs ~ pub mod simulation
│ │ ├── simulation.rs + RunSetup, Simulation, StatSample, SetupError, ExportError
│ │ ├── entry/mod.rs ~ setup, schema
│ │ └── gpu/stepping.rs ~ sample_stats over encode_stats_passes
│ ├── henad-models/src/tests/
│ │ ├── mod.rs ~ mod simulation
│ │ ├── registry.rs ~ every_default_setup_passes_the_checks_of_from_parts
│ │ └── simulation.rs + the setup and simulation tests
│ ├── henad-explore/src/
│ │ ├── lib.rs ~ pub mod benchmark
│ │ ├── benchmark.rs + run_benchmark and its tests
│ │ ├── sweep.rs ~ run_spec, plan_spec, sweep_device, SweepOptions::new, apply_execution
│ │ ├── search_run.rs ~ run_search and NotASweep gone, the folder from the inputs
│ │ ├── handle.rs ~ SweepRunOptions::new, spec_source, plan_for_start, BoundModel gone
│ │ ├── pumped.rs ~ options through new, no panic hook
│ │ ├── spec_file.rs ~ LoadedSpec, SpecFileError::Write
│ │ ├── result_set.rs ~ ModelSchema in place of &ModelEntry
│ │ ├── schema.rs ~ model_schema gone
│ │ ├── probe.rs, exec/mod.rs ~ pub(crate)
│ │ ├── output/details.rs + params_by_id_json, scheduled_actions_json
│ │ ├── output/{manifest,memory,mod,read,runs_csv,series_csv,summary_csv,search_tables}.rs ~ pub(crate)
│ │ └── tests/ ~ support.rs, replay.rs, search.rs, gpu.rs, resume.rs and the other call sites
│ ├── henad-cli/src/
│ │ ├── main.rs ~ benchmark through run_benchmark, exports through Simulation
│ │ ├── actions.rs - moved into henad-explore's benchmark.rs
│ │ ├── explore.rs ~ LoadedSpec, run_spec and plan_spec, flags after the table
│ │ └── json_report.rs ~ the shared JSON helpers
│ └── henad-app/src/
│ ├── state.rs ~ build_setup, setup_message, open_run through from_replay
│ ├── ui/playback.rs ~ a SetupError as Build's disabled reason
│ └── ui/{sweep,results,export}/ ~ SweepRunOptions::new, LoadedSpec::to_toml, entry.schema(), details helpers
└── docs/
├── reference/cli.md ~ the exports through Simulation
└── developing/agent-record/20261001-35-library-api.md +
State after¶
M6 is implemented and uncommitted on 48-library, on top of abd6db1.
The workspace version stays 0.2.0.
HENAD_REQUIRE_GPU=1 ./check.shpasses: 1006 tests, none failed, with the wasm32 typecheck, the packaging, cargo-deny and docs steps and the web build. M5 passed at 991, and M6 before its review at 1002. The first run before the review stopped at rustdoc, on public docs linking to the now crate-privatechoose_layoutand on a link fromSweepRunOptionstoSweepOptionsoutside its scope.- The review's changes touch no measured path.
run_benchmark,run_forandrun_sampledare as the gates timed them, and the gates were not rerun. uv run --locked zensical buildpasses with its path checks.- Every M6 gate of [7.1] passes against the M5 binary, with
run_sampledchanged after its miss and virus_network passing on its rerun. --export-statsand--exportwrite the M5 binary's bytes, and the 0.2.0 outbreak export its bytes.
Proposed commit message: feat: programmatic API
Issues found & future directions¶
run_sampled's callback isSend. [4.2.2] wanted it called outside the pool with noSendbound, at one pool entry per sample. That shape missed the gate by 7.6 times at Game of Life 64², and the bound is the cost of one pool entry per call. A host whose callback holds something that is notSendsteps throughrun_toandstatsin its own loop. The design's [4.2.2] now says so, andguide/library.mdtakes the same wording when M11 writes it.- The app cannot show a refused value on its own. Every Parameters widget clamps, so Build's new reason appears only for fields set from elsewhere. Build checks them all the same, and a future loader that sets the fields directly meets the check.
- A loop of
step()frommaincosts a pool entry per tick. Shape (iv) is 12 times shape (i) at Game of Life 64² and 1.07 times at 4096².Simulation::step's docs say to wrap such a loop inrayon::scopeorpool.install, and shape (iii) shows that this recovers shape (i). - The GPU tick-0 actions fire in a scope of their own after the factory's, with
BUILDINGas theirduring. Moving them into the factory's scope would need the schedule inside the factory. - Reduced rungs. bench_matrix's 100,000 GPU steps would hold gpu_boids at 1M agents for over 9 minutes a repetition, past the cap. The GPU rungs that ran with fewer steps are named in the instrument table.
- Next. M7 (provenance) depends on M3 and M5 and can start. M8 and M9 wait for M7.