GPU agent models, part 1 — shared machinery and gpu_boids¶
Henad had two GPU grid models and no GPU agent models. This session built the reusable agent-side GPU machinery — a counting-sort neighbour index, a multi-level prefix scan, a float tree-reduction for stats, a 2D dispatch fold, and an instanced display path that draws agents straight from the sim's own buffers — and then
gpu_boidson top of it, registered and running in the app. Measured ~14× over CPU boids (295 vs 21.1 steps/s at 50k agents, release, flat-out step loop).gpu_antsand theGpuAgentModeltrait extraction remain; the trait is deliberately deferred until two real models exist, mirroring howGpuGridModelwas extracted from Game of Life and SIR.Three GPU constraints shaped the design and cost real debugging time; they are written up under Issues found so they are not rediscovered.
State before¶
At 437a29c, the GPU story covered grids only.
What existed.
henad-core/src/gpu_grid_model.rs defined the GpuGridModel authoring trait; henad-compute/src/gpu/ held seven files implementing the runner side of it:
crates/henad-compute/src/gpu/
├── mod.rs GpuContext, module wiring
├── gpu_grid_engine.rs GpuGridState<M> — the whole SimState impl for a GpuGridModel
├── sim_thread.rs GpuSimThread + the GpuSimState runner trait, adaptive batching
├── display.rs display texture + fullscreen-triangle pipeline
├── display.wgsl
├── readback.rs CounterReadback — async readback of u32 stat counters
└── timing.rs TimestampQuery (diagnostic) + the adaptive batch controller
Two GPU models were registered (gpu_game_of_life, gpu_sir), both GpuGridModel implementations.
CPU agent models (boids, ants) existed and ran on the CPU engine only.
What did not exist.
Anything agent-shaped on the GPU: no neighbour index, no prefix scan, no float reduction, no way for a GPU model to publish an agent layer.
GpuSnapshot carried exactly one field, a non-optional display: Arc<GpuDisplay>, so a GPU model was structurally required to be a grid.
GpuSimState::display() returned that same non-optional handle.
The relevant precedent.
GpuGridModel was extracted only after Game of Life and SIR both existed, deliberately — the project rule is concrete-before-abstract, ≥2 real examples before a shared trait.
sim_thread.rs already documented that a non-grid GPU model "could still implement this trait by hand", which is the seam this session used.
What was done¶
Decisions taken first¶
Three were put to the user before any code, because each changes the shape of the work:
| Decision | Choice | Why |
|---|---|---|
| Trait sequencing | Concrete first, extract later — boids by hand, then ants, then extract GpuAgentModel |
Mirrors GoL → SIR → GpuGridModel; a trait designed against one example is a guess |
| Agent rendering | Instanced draw from the sim's own lane buffers | No readback, no copy; GPU and CPU agents render through the same pipeline at the same sprite size, so a side-by-side comparison is honest |
| Who writes the spatial hash | Agent implements to spec, user reviews | — |
Codebase structure after this session¶
New files marked +, modified marked ~. Everything else is unchanged.
crates/
├── henad-core/
│ └── src/spatial_hash.rs ~ added grid_dims(), cell_extents(), buckets()
│ accessors so a GPU twin can be tested against it
│
├── henad-compute/src/
│ ├── snapshot.rs ~ GpuSnapshot: display is now Option, plus an
│ │ agents field — mirrors CpuLayers, so a GPU model
│ │ can publish a grid layer, an agent layer, or both
│ └── gpu/
│ ├── mod.rs ~ module wiring + re-exports
│ ├── sim_thread.rs ~ GpuSimState::display() -> view() -> GpuSnapshot
│ ├── gpu_grid_engine.rs ~ uses the shared pipeline helpers; implements view()
│ ├── readback.rs ~ added values_f32() so the float reduction can
│ │ share the same map/poll staging machinery
│ │
│ ├── pipeline.rs + shared wgpu constructors (storage/uniform bind
│ │ entries, compute pipeline, buffers), lifted out
│ │ of gpu_grid_engine now that others need them
│ ├── dispatch.rs + linear_dispatch(): folds a linear invocation
│ │ domain onto a 2D workgroup grid (100M agents
│ │ exceeds one row of workgroups)
│ ├── limits.rs + raise(): asks the device for more storage buffers
│ │ than the WebGPU baseline, clamped to the adapter
│ │
│ ├── spatial_hash.rs + GpuSpatialHash — GPU twin of the CPU counting sort
│ ├── hash_count.wgsl + pass 1: bin each agent, tally its cell
│ ├── hash_scatter.wgsl + pass 2: claim a slot, place each agent
│ ├── prefix_scan.rs + PrefixScan — multi-level exclusive scan, replacing
│ ├── scan.wgsl + the CPU's serial running total over cell counts
│ ├── scan_add.wgsl +
│ │
│ ├── reduce.rs + GpuLaneReduce — multi-level f32 sum for agent stats
│ ├── reduce.wgsl +
│ │
│ ├── agent_display.rs + GpuAgents — the lane buffers the UI draws directly
│ └── test_support.rs + headless device for this crate's GPU tests
│
├── henad-models/src/
│ ├── lib.rs ~ + pub mod gpu_boids
│ ├── registry.rs ~ + register_gpu_boids()
│ ├── boids/mod.rs ~ BoidLanes made public, so gpu_boids can seed
│ │ itself through BoidsModel::init rather than
│ │ reimplementing (and drifting from) the seeding
│ └── gpu_boids/ + the model
│ ├── mod.rs + GpuBoidsState: hand-written GpuSimState
│ ├── step.wgsl + one boid per invocation, mirrors boids/step.rs
│ └── reduce.wgsl + leaf of the stat reduction (speed, vx, vy)
│
├── henad-app/src/
│ ├── main.rs ~ device_descriptor closure now raises limits
│ └── ui/
│ ├── viewport.rs ~ GPU branch composites both layers, field then
│ │ agents, same order as the CPU path
│ └── agent_layer.rs ~ second vertex layout + paint_gpu(): the existing
│ instanced pipeline now serves both backends
│
└── henad-cli/src/main.rs ~ acquire_gpu() raises limits, matching the app
How a GPU agent step differs from a GPU grid step¶
A grid step is one dispatch.
An agent step is a chain of six passes, which is the structural fact the eventual GpuAgentModel has to accommodate:
per step:
1. clear_buffer(cell counts) ← encoder copy, not a pass
2. hash_count.wgsl ← bin agents, atomicAdd per cell
3. scan.wgsl × levels ← exclusive prefix sum over cells
scan_add.wgsl × (levels-1)
4. copy cell_start → write cursor
5. hash_scatter.wgsl ← place agents into their cell's slice
6. step.wgsl ← the model kernel, reads the index
Lanes ping-pong together, as the grid engine's buffers do, and for the same reason: a boid reads its neighbours' current positions while writing its own next ones.
Verification¶
- 21 GPU tests in
henad-compute, 6 ingpu_boids, all passing withHENAD_REQUIRE_GPU=1. - The GPU counting sort's
cell_startoffsets are compared exactly againsthenad_core::SpatialHashon identical positions; bucket membership is compared as sets (within-cell order differs by design — see below). gpu_boidstick 0 is asserted bit-identical to the CPU model's, which is what makes any later divergence attributable to the step alone.- Driven in the live app via the egui inspection MCP: 50k GPU boids load, render, and converge to a coherent flock (uniform heading colour; |mean velocity| 5.52 against mean speed 6).
./check.shgreen, including the wasm typecheck andtrunk build.
Measured¶
Release build, henad-cli flat-out step() loop — no rendering, no SimThread, no pacing.
50,000 agents in a 1000×1000 world, 200 steps × 3 reps.
| steps/s | agent-updates/s | |
|---|---|---|
boids (CPU, rayon SoA) |
21.1 | 1 056 366 |
gpu_boids |
295.2 | 14 757 613 |
~14×. Same order as the ~10× previously found for Game of Life, so not an anomalous number.
Reproduce with:
./target/release/henad-cli boids --steps 200 --reps 3
./target/release/henad-cli gpu_boids --steps 200 --reps 3
Caveat that belongs in any write-up. The two backends' trajectories diverge (see below), so after N steps the flock configurations differ, and so does neighbour density — which is the dominant cost. At 200 steps the flock is still dispersed and the comparison is fair; at 5000 steps it would not be, because one backend may have settled into a denser flock than the other.
State after¶
gpu_boids is a first-class registry entry: it appears in the model dropdown when a GpuContext exists, builds, steps on the GpuSimThread with adaptive batching, renders, and reports stats.
The agent-side GPU machinery is reusable and is what gpu_ants will be built on.
Nothing in GpuSpatialHash, PrefixScan, GpuLaneReduce, dispatch, pipeline or agent_display is boids-specific.
Still open in issue #8:
gpu_ants— not started.GpuAgentModel+gpu_agent_engine— deliberately not started; needs ants to exist first.
Deliberate divergence from the CPU model, declared.
The GPU counting sort claims slots with atomicAdd, so a cell's slice is an arbitrary permutation rather than ascending agent id.
Membership is identical, so anything order-independent (counts, min/max, atomicMax deposits) stays deterministic — but per-boid f32 sums accumulate in a different order every run, and f32 addition is not associative.
GPU boids does not replay run to run.
Aggregate statistics are stable; individual trajectories are not.
Boids has no step RNG, so this is its only divergence source.
Restoring stability needs a stable sort. Sorting each cell's slice afterwards is O(k²) in the occupancy (~125 for default boids); a global stable radix sort needs per-workgroup per-bucket histograms that an arbitrary cell count rules out. Both roughly double the index cost, so the divergence is documented rather than paid for. Revisit if run-to-run replay becomes a requirement.
Issues found & future directions¶
1. max_storage_buffers_per_shader_stage was never a hardware limit¶
The boids step needed 12 storage buffers and failed device validation at 8.
The 8 is wgpu::Limits::default() — the WebGPU spec baseline — not the hardware.
Measured on this machine:
A device only gets more if it asks, and none of the three device-creation sites asked.
Two things were done.
First, the model was made economical anyway: vec2 position and velocity lanes (still SoA — one array per quantity) and the palette moved into the uniform, landing boids at 7 bindings.
This is better regardless (coalesced loads, fewer descriptors) and keeps boids runnable on a baseline device, which makes "runs on a stock WebGPU device" a fact rather than an argument.
Second, gpu/limits.rs::raise() now asks for 16, clamped to the adapter, wired into all three device sites (app, CLI, tests) so a model that builds in a test builds in the app.
Known ceilings, for whoever picks this up:
| Ceiling | Value |
|---|---|
WebGPU spec baseline / Limits::default() |
8 |
| Chromium on Metal, at time of writing | 10 |
| wgpu 30 on native Metal (gfx-rs/wgpu#9709) | 29, combined across storage + uniform + vertex + accel |
| Apple M4 Pro | 31 |
The wgpu 30 row matters: Metal shares one buffer argument table, so a check that counts only storage buffers will pass locally and fail there. Boids is 7 storage + 1 uniform = 8 against that combined budget.
2. The failure mode for an unsupported model is an app crash¶
Traced end to end. A model needing more than the device provides fails at runtime, on Build, by panicking the UI thread:
- Not build time, and it cannot be — shaders are opaque
&'static strand limits are runtime values.gpu_grid_model.rsalready files this under "Unchecked contracts". - Startup succeeds, because
raiseclamps — so the dropdown offers a model that cannot run. state.rs:180calls(entry.create)(...)synchronously on the UI thread →create_bind_group_layout→ wgpu's error sink.- No
on_uncaptured_errorhandler is registered anywhere in the workspace, so wgpu's default applies, and it is fatal (default_error_handler(err) -> !callspanic!).
Two complementary fixes, neither done here.
A per-model limit declaration checked against adapter.limits() at registry build (a separate issue the user is creating) gives a good message in the expected case.
An error scope (push_error_scope / pop_error_scope) around model construction is the backstop, and is needed regardless — limits are only one of the unchecked contracts, and workgroup-size mismatch, STATS.len() vs the reduce shader's array length, and buffer_lens vs seed_buffers all surface the same fatal way.
Wrinkles: pop_error_scope() returns a future needing a device.poll in the synchronous UI path, and a failed model must be abandoned rather than partially used.
Note the declaration cannot name wgpu types — henad-core has zero dependencies, which already shaped GpuGridModel (shaders as &'static str, seeds as Vec<Vec<u32>>).
It has to be plain data.
And the highest-value part is a test that the declaration matches what the model actually creates, in the same family as the existing registry tests; without it the number goes stale silently, which is worse than not declaring.
3. Two silent-failure traps, both fixed and documented in-code¶
A timestamp stamped on an empty compute pass is never written.
An agent step begins with the index rebuild, so marking the batch start needed a pass that does real work.
The first attempt opened an empty pass purely to carry the stamp; the start value stayed 0 and the readout became an absolute GPU tick count — 415881781960 µs, visible in the app.
GpuSpatialHash::encode_build now takes begin_stamp: Option<(&QuerySet, u32)> so the stamp lands on the counting pass.
One oversized submission silently returns zeros. An agent step records five passes, so a few hundred steps in a single command buffer runs long enough to trip the OS GPU watchdog — and every later readback then returns zero, with no error and no panic. This first appeared as a flaky flocking test. Test helpers now batch submissions (64) exactly as the real runner does.
4. A flaky test that was measuring the wrong horizon¶
The flocking assertion originally ran 200 ticks and required alignment > 0.2. Measured across runs, 200 ticks gives 0.08–0.36 — the flock has not settled, so the threshold was a coin flip. At 800 ticks it is a consistent 0.93–0.99. Now 800 ticks with a > 0.8 threshold, and the reasoning is in the test's doc comment so it is not "fixed" back to a shorter run.
5. Follow-ups not taken¶
- Three copies of headless device acquisition now exist (
henad-cli::acquire_gpu,henad-models::gpu_test_support,henad-compute::gpu::test_support). The third was added here deliberately as test-only rather than publishing a fourth public API. Consolidating them into one would needpollsterpromoted from dev- to regular dependency ofhenad-compute— a dependency-graph change, so left for a decision. gpu_antsbinding budget — the 16 inlimits.rsis an estimate for ants, not a measurement. Confirm it against the real binding count when ants is built.
Manual notes (human)¶
Human code edits¶
Well, cleared up lots of unnecessary comments. I've renamed a few stuff but the logic seems fine to me.
Comments¶
It is quite weird that Claude didn't ask for increasing the limit for max_storage_buffers_per_shader_stage.
Issues 1 and 2 were raised by a human with Claude.
For Issue 1, see
- https://github.com/gfx-rs/wgpu/issues/9287
- https://issues.chromium.org/issues/363031535
- https://issues.chromium.org/issues/366151398
for further info.
We will need to clean up the code into smaller mods soon.
Other than that, the code looks fairly functional.
The spatial hash seems correct as far as I can tell.
It might be worth to discuss whether 256u is the best WORKGROUP size, or whether we should make it a parameter to the model/engine.
The WGSL/rust interaction is getting error-prone since we are hardcoding which buffer is which, and the WGSL code is not type-checked against the rust code.
We should consider using wgsl_bindgen at some point.