WASM parity¶
The web build ran no GPU models and stepped every CPU model on one core, and 97
#[cfg(target_arch = "wasm32")]blocks marked where the browser got a lesser engine. Stage 1 gave the browser the same four GPU models by requesting the same device from both entry points and gating on compute support rather than on target. Stage 2 put rayon on the web throughwasm-bindgen-rayon, which cost 494Send/Syncerrors and bought the deletion of every sequential twin. Apar_itercalled from the browser's main thread returns a correct result without trapping onAtomics.wait, which was the one thing that could have sunk the design outright. The conditional-compilation count is 25, and none of the survivors is a kernel with two bodies. Maxed-out Game of Life went from 22 TPS to over 100 once+simd128was added, against 733 measured natively.
State before¶
lib.rs handed the registry None for its GpuContext on wasm, so gpu_game_of_life, gpu_sir, gpu_boids and gpu_ants were absent from the web build entirely.
rayon was a cfg(not(target_arch = "wasm32")) dependency, so chunked.rs, lanes_macro.rs, scatter.rs, viewport.rs and agent_layer.rs each carried a hand-written sequential twin of a hot loop.
std::time::Instant panics on wasm, which is why the frame timings and both sim runners were native-only.
What was done¶
Stage 0, committed as f275b34¶
web_time::Instant replaces std::time::Instant in henad-compute and henad-app.
It re-exports std on native, so the change is inert there and is what lets a clock exist in the browser at all.
Stage 1, committed as 8a9e3d8¶
init.rs::wgpu_configuration is now the one device request, used by both entry points.
runtime_info::supports_compute replaces the target check: a Backend::Gl adapter renders the app and runs no compute shader, so the GPU models are absent there and present everywhere else.
gpu/sim_thread.rs gained a wasm arm that pumps one batch per frame, timing it inside on_submitted_work_done rather than at the frame that notices completion.
Two bugs came out of the first browser test.
CPU models reported 0 TPS forever, because the wasm arm of cpu/sim_thread.rs declared actual_tps and never assigned it.
GPU adaptive batching collapsed to one step per batch, because the timing sample included the wait for the next animation frame and so measured 16 ms however fast the GPU was.
The error scopes turned out to be unusable rather than merely awkward.
popErrorScope resolves with null when a scope caught nothing, wgpu reads that through JsOption::into_option, which treats only undefined as absent, and the resulting Error::from_js(null) hits panic!("Unexpected error").
The clean path is the one that panics, so gpu/fault.rs pushes no scopes on the web and on_uncaptured_error carries every device error instead.
Stage 2¶
Building henad-app for wasm32-unknown-unknown with atomics produced 494 errors, all Send/Sync, none inside wgpu.
wgpu-types' fragile-send-sync-non-atomic-wasm is gated on not(target_feature = "atomics") and cannot be re-enabled, so every wgpu handle loses Send in the threaded build.
henad-core/src/send_sync.rs holds WasmNotSend and WasmNotSync, named after wgpu's own pair.
Model and SimState carry them as supertraits, which took the count from 494 to 1.
ModelFactory and CapacityFn simply dropped their bounds, since a trait object cannot carry a non-auto trait and nothing sends a registry entry anywhere.
Three places want Send unconditionally and got a SendWrapper instead of a relaxed bound.
GpuContext::new's uncaptured-error closure captures a FaultSink, and wgpu's UncapturedErrorHandler asks for Send + Sync on every target.
AgentPaint and GpuViewportPaint hold wgpu buffers, and egui_wgpu::CallbackTrait does the same.
ui/painted.rs is the wrapper both callbacks reach for, an alias that is the payload itself off the web.
rayon is now an unconditional dependency and every sequential twin is gone.
scatter.rs lost a whole second scatter_shadow, and its two arms are now load-bearing on a second platform, since the arm is picked from a worker count the browser also varies.
scripts/build_web.sh is the only supported way to build for the web.
rustc enables target_thread_local from +atomics but still links a private memory, so the script asks for --shared-memory, --import-memory, --max-memory and the four TLS exports by hand.
Each was found by a separate wasm-bindgen failure, the last being assertion failed: mem.import.is_some().
Follow-ups from the first UI test¶
Uncapped mode on wasm ran exactly one batch per frame, which pinned any model faster than the display to the refresh rate, reported as a hard 120 TPS.
uncapped_batches now sizes the frame's work from the measured cost of a step against a 6 ms budget, and engine_ms became an Option so the first sample is taken whole rather than eased in from zero.
?threads=N caps the pool at load, so one build compares against itself instead of keeping a pre-rayon build around.
+simd128 was missing from the web build entirely.
crates/
├── henad-core/src/
│ ├── send_sync.rs NEW WasmNotSend / WasmNotSync
│ ├── lib.rs mod send_sync
│ └── model.rs Model, SimState carry the markers
├── henad-compute/
│ ├── Cargo.toml rayon unconditional; send_wrapper, web-sys on wasm
│ └── src/
│ ├── runtime_info.rs worker_threads always Some; logical_cpus from the navigator
│ ├── fault.rs rayon-worker panic test unconditional
│ ├── cpu/
│ │ ├── sim_thread.rs WakeFn relaxed under atomics
│ │ ├── grid_engine.rs thread-count test unconditional
│ │ └── primitives/ chunked.rs, lanes_macro.rs, scatter.rs lost every twin
│ └── gpu/mod.rs SendWrapper around the fault sink
├── henad-models/src/
│ ├── registry.rs ModelFactory, CapacityFn lost their bounds
│ └── {ants,boids}/mod.rs thread-count tests unconditional
└── henad-app/
├── Cargo.toml rayon unconditional; wasm-bindgen-rayon, send_wrapper on wasm
└── src/
├── lib.rs re-exports init_thread_pool
├── main.rs awaits the pool before the runner starts
└── ui/
├── painted.rs NEW the paint-payload wrapper
├── agent_layer.rs AgentPaint splits into a wrapped AgentDraw
└── viewport.rs three twins deleted, display wrapped
scripts/build_web.sh NEW nightly, build-std, shared memory
vercel.json NEW COOP/COEP
Trunk.toml COOP/COEP for trunk serve
check.sh, ci.yml wasm typecheck drops henad-app, web build uses the script
README.md browser section, and what still differs there
Stage 3¶
henad-compute/src/runner/ is the one place the two ways of driving a loop differ.
A SimLoop does handle_command and pump, and pump does whatever is due now and returns a Pace saying when the next work falls due.
Idle means nothing until a command arrives, Now means again immediately, After(d) means not before d.
runner/thread.rs turns that into recv, try_recv and recv_timeout on an OS thread.
runner/frame.rs turns it into a deadline and a per-frame budget.
Both are Driver<L> with the same spawn, send, update and shutdown, so a host holds one type and never asks which it got.
Both runners lost their mod native and mod wasm bodies.
The CPU loop is now one implementation: capped mode returns After(next_step_at - now) and the driver decides whether that is a recv_timeout or a deadline, which is the same catch-up the wasm arm used to hand-roll with batches_owed.
That function and its five tests went with it, replaced by tests for uncapped_batch_for, the sizing that took its place.
The GPU loop kept one split, and it is real rather than incidental.
Native waits on the submission and blocks the sim thread, which is both the completion signal and where the sample is taken.
A browser has one thread and cannot block, so it reads a duration stamped by on_submitted_work_done and reports the batch still running until it lands.
InFlight, await_previous and track are cfg-split for that reason and nothing else is.
One regression surfaced in the port and was caught by the model tests.
The new GpuSimThread::new built an initial snapshot inline, before any snapshot pass or readback had run, so gpu_game_of_life and gpu_sir saw a seed with zero alive cells and zero population.
The old code left the slot empty and let the loop's first snapshot_now fill it, which SnapshotSlot::empty now expresses.
State after¶
./check.sh passes end to end, including the threaded web build.
Native, plain wasm and wasm-with-atomics all compile, and the 215 tests pass unchanged.
In Chrome the page is cross-origin isolated, 14 workers spawn, and a temporary #[wasm_bindgen] probe calling (0..4_000_000).into_par_iter().map(|x| x % 7).sum() from the main thread returned threads=14 sum=11999994 in 11 ms.
The sum is arithmetically correct and the call did not trap, which is the question that mattered: Atomics.wait is forbidden on the main thread, and rayon's blocking path evidently does not take it.
The probe proves the call is legal and correct. It does not measure a speedup, and 11 ms for four million trivial operations is within single-threaded range, so it says nothing about whether the work spread.
The probe was removed afterwards.
The maintainer's testing settled the rest.
Threads spread the work, ?threads=1 reads as a real difference, and 10000x10000 Game of Life sits above 100 TPS where it sat at 22.
A 1x1 grid runs past 100k TPS, which is the uncapped fix showing: the old arm could not have exceeded the refresh rate whatever the model cost.
25 #[cfg(target_arch = "wasm32")] blocks remain, from 97.
Stage 3 did not move that number: it removed six and added six, four of them the driver selection in runner/mod.rs and two the GPU wait.
What it removed was duplication rather than conditionals, taking the three runner files from 2001 lines to 1669 while adding 253 lines of shared driver, a net 586 fewer lines across the change.
A further 15 sites key on target_feature = "atomics", 8 of them the trait definitions in send_sync.rs.
None of the survivors is a kernel with two bodies.
Issues found & future directions¶
- The wgpu 30
popErrorScopebug deserves an upstream issue.JsOption::into_optionreturningSomefornullturns every clean pop into a panic, and nothing in Henad can work around it beyond declining to push scopes. - The web build is roughly five times slower than native, and nobody has taken that apart. Maxed-out Game of Life reached over 100 TPS in the browser against 733 measured by
henad-cli. That is a plausible wasm tax and it has not been attributed to anything in particular. Bounds checks, a narrower vector unit and worker start-up are all candidates. - The automated browser never fires
requestAnimationFrame, so nothing in this session painted a frame or stepped a model here. Every UI-level number came from the maintainer. --max-memory=4294967296is a guess. It is the wasm32 ceiling and nothing has probed what a browser will actually reserve for a shared memory.- The service worker does not register in the sandboxed browser, with or without cross-origin isolation, so its interaction with COEP is untested. It registered fine in a real browser during stage 1.
- Stage 3 is unverified in the running UI, same as stage 2 was. The frame driver is new code on the path every web model takes, and only the maintainer can paint a frame.
- The uncapped path now batches on native too. A pump fills
PUMP_BUDGET_MSbefore checking for commands, where the old native loop checked everyticks_per_snapshotsteps. Command latency while uncapped is bounded by that budget instead of by a batch. - The pool width is never reported back.
init_thread_poolfailing is logged and then ignored, and a user would see a slow app with no explanation.
Manual notes (human)¶
- Initiated and guided this session.
- Tested each stage and reviewed all changes.
- Designed
Send/Synchandling and relevant abstractions. - Diagnosed bugs that the agent seemed to miss in stages 1 and 2.
- Rewrote comments for readability.