GPU agent models, part 2 — gpu_ants¶
The second GPU agent model, and the one that decides what the eventual
GpuAgentModeltrait has to cover.gpu_antsis a composite model — it publishes both snapshot layers, a pheromone texture and an agent layer — and it is the first GPU model with a field that agents write. Measured 7.7× over CPU ants at the default 2k/200², rising to 18.4× at 200k/1000².Two things fell out of the port that are worth carrying into the trait design. Ants needs fewer passes than boids, not more: two, against boids' five, because it has no neighbour index and no ping-pong. And it lands on exactly eight storage buffers, the WebGPU baseline, which the tests prove rather than argue —
gpu_test_supportasks forLimits::default().Unlike
gpu_boids,gpu_antsreplays bit-identically run to run, and there is a test asserting it.
State before¶
At add53cd, on branch 8-gpu-agent-models.
What existed.
The agent-side GPU machinery from the previous session — GpuSpatialHash, PrefixScan, GpuLaneReduce, linear_dispatch, the pipeline constructors, GpuAgents — plus gpu_boids built on it.
GpuSnapshot already carried an optional display layer and an optional agent layer, and the viewport already composited both.
Three GPU models were registered: gpu_game_of_life, gpu_sir, gpu_boids.
What did not exist.
Any GPU model with a field the agents write, and therefore any use of the scatter-plus-decay pattern on the GPU.
No GPU model published both snapshot layers at once — the two grid models published a texture, gpu_boids published lanes.
The relevant precedent.
gpu_sir established the per-cell self-advancing RNG buffer, and its step.wgsl documents why: a uniform written between passes in one encoder only becomes visible after the whole encoder submits, so every pass in a batch would read the same tick.
gpu_boids established seeding through the CPU model's own init, so tick 0 is bit-identical and any later divergence is attributable to the step.
What was done¶
Codebase structure after this session¶
New files marked +, modified marked ~.
crates/henad-models/src/
├── lib.rs ~ + pub mod gpu_ants
├── registry.rs ~ + register_gpu_ants(), topology grid+agents
├── ants/mod.rs ~ AntLanes and NO_STEP re-exported, so gpu_ants can seed itself
│ through AntsModel::init rather than reimplementing the seeding
└── gpu_ants/ + the model
├── mod.rs + GpuAntsState: hand-written GpuSimState
├── step.wgsl + one ant per invocation, deposit and advect fused
├── merge.wgsl + one cell per invocation, max-combine then decay
├── display.wgsl + quantise the field into the display texture
└── reduce.wgsl + leaf of the stat reduction (carrying, total pheromone)
Nothing in henad-core or henad-compute changed.
The machinery from part 1 was sufficient, which is itself a useful signal for the trait extraction.
A GPU ants step is two passes, not five¶
This is the structural surprise, and it is the opposite direction from what was expected.
gpu_boids, per step: gpu_ants, per step:
1. clear cell counts 1. step.wgsl — deposit + advect
2. hash_count.wgsl 2. merge.wgsl — combine + decay
3. scan / scan_add
4. copy cell_start
5. hash_scatter.wgsl
6. step.wgsl
Three things collapse:
- No neighbour index. Ants use
NoIndexon the CPU for the same reason — they never read one another, only the field. The whole hash/scan/scatter chain is absent. - No ping-pong on the lanes. Every
AntLaneslane isplain, so each invocation reads and writes only its own slot. - No ping-pong on the field. Deposits land in a separate
accumbuffer, sofieldis read by pass 1 and updated in place by pass 2.merge.wgslalso zeroesaccumas it consumes it, which is free — the pass already owns the cell — and removes the per-tickclear_bufferthe design started with.
Fusing the CPU's two agent passes is exact, not an approximation.
run_deposit_pass and run_step_pass both read the field as it stands before ScalarField::update, and the only lanes the deposit reads (reward) are ones the advect later writes for the same agent.
One invocation doing both in order is the same computation.
Binding budget: exactly eight storage buffers¶
The part-1 note said "ants will need more than boids, so budget bindings from the start." It needs the same: eight.
| binding | buffer | why it cannot merge with another |
|---|---|---|
| 0 | pos |
vec2<f32>, also the renderer's instance stream |
| 1 | state |
packs last_step, has_food, has_reward |
| 2 | color |
packed RGBA, the renderer's other instance stream |
| 3 | rng |
per-agent self-advancing state |
| 4 | field |
both pheromone layers, layer f at f * n_cells |
| 5 | accum |
this tick's deposits, atomic<u32> |
| 6 | sites |
static terrain |
| 7 | counters |
the cumulative delivery count |
What gets it to eight is reward being one bit rather than an f32 lane.
The CPU model only ever stores 0.0 or the reward param in that lane — advect_agent sets it to 0.0 and only a site grants p.reward.
the_cpu_reward_lane_only_ever_holds_two_values asserts that over 200 ticks of the CPU model, so the packing cannot rot silently.
Since gpu_test_support requests Limits::default(), the passing test suite is direct evidence that gpu_ants runs on a stock WebGPU device.
Against the Metal combined argument-table budget of 29 it is 8 storage + 1 uniform = 9.
If a ninth is ever needed, the escape hatch is packing state and rng as one vec2<u32> lane — they are read and written by the same invocation, so interleaving them would if anything help locality.
The deposit combine¶
Combine::Max maps onto atomicMax on the u32 bit pattern of the deposit, which is valid because every deposit is non-negative (field values, cutdown, and reward all are) and IEEE-754 ordering matches unsigned integer ordering there.
This is the part that makes ants deterministic where boids is not.
Boids diverges run to run because the counting sort's within-cell order is arbitrary and f32 addition is not associative; max is order-independent, so no such thing happens here.
a_run_replays_bit_identically compares positions, packed state and the whole field between two fresh runs of 300 ticks.
Verification¶
- 9
gpu_antstests, all passing withHENAD_REQUIRE_GPU=1, alongside the existing 47. - Tick 0 positions asserted bit-identical to
AgentModelState::<AntsModel>. - The behavioural test is
the_colony_lays_a_trail_and_delivers_food: after 1500 ticks, pheromone is non-zero, ants are carrying, and deliveries have happened. None of that is true if the deposit never lands, the merge never runs, or the trail is followed in the wrong direction. total_pheromone_agrees_with_the_fieldcross-checks the two-lane reduction against a readback of the field it summed.- Driven in the live app via the egui MCP: at tick 22k the colony has consolidated a tight two-lane highway around both obstacle blobs, 80 833 deliveries, both gradients rendering.
./check.shgreen, including the wasm typecheck andtrunk build.
Measured¶
Release build, henad-cli flat-out step() loop — no rendering, no SimThread, no pacing.
300 steps × 3 reps, --global-warmup 100, CPU and GPU runs interleaved and repeated to control for this machine's thermal drift. Both passes agreed to within 1.5%.
| agents / world | ants (CPU) steps/s |
gpu_ants steps/s |
speedup |
|---|---|---|---|
| 2 000 / 200² (defaults) | 1 468 | 11 330 | 7.7× |
| 50 000 / 500² | 648 | 9 285 | 14.3× |
| 200 000 / 1000² | 233 | 4 295 | 18.4× |
Reproduce with:
./target/release/henad-cli ants --set num_agents=50000 --set world_width=500 --set world_height=500 --steps 300 --reps 3 --global-warmup 100
./target/release/henad-cli gpu_ants --set num_agents=50000 --set world_width=500 --set world_height=500 --steps 300 --reps 3 --global-warmup 100
Why the default config understates it. At 2k ants over 200², the merge pass runs 80 000 invocations against the step pass's 2 000, so the step is entirely cell-bound and the GPU is mostly idle in the pass that matters. The ratio climbs with agent density, which is why the 200k row is the honest headline for "how much does the GPU buy you" and the 2k row is the honest headline for "what does the shipped default do".
Caveat that belongs in any write-up. The RNG streams differ (below), so the two backends' colonies take different paths and reach different trail densities. Trail density is the dominant cost, since it decides how often the tie-break draw fires. The runs above are short enough that both are still in the exploratory regime; a 20k-tick comparison would not be fair without checking that both had settled.
State after¶
gpu_ants is a first-class registry entry: it appears in the dropdown when a GpuContext exists, builds, steps on the GpuSimThread with adaptive batching, renders both layers, and reports all three stats.
Issue #8 now has both models. The remaining work is the GpuAgentModel + gpu_agent_engine extraction, which was deliberately deferred until two real models existed. They now do.
What the two models tell the trait design¶
Worth reading before designing GpuAgentModel, because the two are less alike than expected:
gpu_boids |
gpu_ants |
|
|---|---|---|
| passes per step | 5 | 2 |
| neighbour index | GpuSpatialHash |
none |
| lane ping-pong | yes, all lanes | none |
| field | none | read by the step, updated by a second pass |
| snapshot layers | agents | display + agents |
| stats | one float reduce | one float reduce over a mixed domain + a persistent counter |
| replays run to run | no | yes |
So the trait cannot assume a pass count, an index, or ping-pong. The honest shape is closer to "a model declares its lanes, its optional index, its optional field, and a list of kernels" than to the grid engine's fixed step/display/reduce triple.
Declared divergences from the CPU model¶
- RNG stream. Per-agent self-advancing
pcg_hash, not per-chunk-per-tickxorshift64. Same reason asgpu_sir. The CPU'snext_unitis(rng >> 40) / 2^24on[0,1); the GPU's isf32(h) / 2^32-1on[0,1]. Different streams either way, so the interval difference costs nothing. rewardas one bit readsparams.rewardat use time, so a live reward edit would differ from the CPU model's ants still carrying the old value. Unreachable in practice: GPU params are reload-only.- Display quantisation uses
log2(v) * 0.30103where the CPU useslog10(v), so a cell sitting exactly on a ramp boundary can land one step either way. Display only.
Everything else is shared by construction: seeding goes through AntsModel::init, sites through PheromoneField::build_sites, both palettes are packed from the CPU model's own arrays, and the hot params come from AntsModel::from_params / PheromoneField::from_params.
Issues found & future directions¶
1. The first TPS sample after Play is garbage, for every GPU model¶
Pre-existing, in henad-compute/src/gpu/sim_thread.rs, not introduced here — found while verifying gpu_ants in the live app, where the Performance panel read TPS 1 523 809 524.
SimCommand::Play resets tps_timer and step_count but not last_stats_publish.
So if the app has been sitting paused for more than STATS_INTERVAL, the very first step_batch after Play sees want_timing == true and calls refresh_tps(now) with now captured microseconds after tps_timer was reset — dividing a whole batch by a near-zero window.
64 steps / ~42 ns is 1.52e9, which is exactly the number observed.
It self-corrects on the next refresh a second later, which is why it is easy to miss: gpu_boids does it too, but at 34 FPS the bogus value is overwritten before you can screenshot it, whereas at 121 FPS gpu_ants still had it on screen.
Left unfixed deliberately — it is shared runner code and not this model's bug.
The one-line fix is for Play to also reset last_stats_publish; a refresh_tps that ignores windows far shorter than STATS_INTERVAL would be the more defensive version.
2. Ants pile into the top-left corner, and it is the reference's bug, not the port's¶
Raised by the user from a viewport screenshot, and present on both backends, which is what made it worth chasing.
ants/step.rs starts the neighbour tie-break's reservoir counter at 2, so with k equally-scented passable neighbours the first one visited gets 2/(k+1) and the rest get 1/(k+1) — 2/9 against 1/9 at the usual k = 8, where uniform sampling would give ⅛.
The visit order is dx outer then dy inner from -1, so the favoured direction is up-left, and momentum 0.8 compounds it into a drift.
Henad is faithful. krABMaga's antsforaging/src/model/ant.rs has the same count = 2 at :115, the same reset at :143, and the same 1. / count at :148.
And gpu_ants matches CPU ants to the same magnitude, so the GPU port is faithful in turn.
Measured with momentum and random action at 0, so the tie-break is the only thing moving ants (2000 ants, 200²):
| mean position | up-left of the food | |
|---|---|---|
CPU, count = 2 |
(16.1, 16.4) | 1719 |
CPU, count = 1 |
(151.4, 151.5) | 0 |
GPU, count = 2 |
(9.2, 9.9) | 1858 |
At default parameters it is a transient, peaking near tick 1000–2000 and gone by tick 10 000 once the trail recruits the strays; count = 1 cuts the peak from 44 strays to 2.
Decision (user, this session): stay faithful and document.
A corrected Henad would stop being the same simulation as krABMaga, and the CPU-vs-CPU comparison depends on it being the same one.
Written up in docs/refs/ants-krabmaga-vs-henad.md under A reference defect Henad reproduces on purpose, with a two-line comment at each of the two tie-break sites so it does not get "fixed" by someone reading it as an off-by-one.
Note for the benchmark work: the bias points at the food, which sits in the same corner, so the defect makes the reference find food faster than a correct model would.
3. The default config is cell-bound, and the parameters cannot express otherwise¶
num_agents and the world extent are independent parameters, so the merge pass's cost (2 × w × h) and the step pass's cost (n) move independently, and the shipped default sits at a 40:1 ratio in favour of the cells.
Nothing is wrong, but it means the default gpu_ants benchmark measures the field update, not the agents.
scripts/bench_matrix.py scales agent models "at constant density", which is the right thing here — worth confirming it picks up gpu_ants and that the density it holds constant is the one that matters.
4. world_width defaults render as 201, not 200¶
Visible in the app's Parameters panel for both ants and gpu_ants.
Cosmetic and pre-existing (the slider's 50.0 step against a 200.0 default), but it means extent.cells() is 201×201 in the GUI and 200×200 from the CLI, so a GUI-vs-CLI number comparison is not comparing the same field size.
5. Follow-ups not taken¶
limits.rs::raise(16)was an estimate for ants in part 1. Ants actually needs 8, so nothing on the current model list needs the raise at all. Worth deciding whether to keep it (headroom for the trait extraction) or drop it back and make "every model runs on the baseline" an invariant with a test.- The WGSL/Rust binding correspondence is still hand-maintained, and this model added four more shader/struct pairs to keep in sync. The part-1 note about
wgsl_bindgenapplies with more force now:StepParamsalone has eleven fields whose offsets are checked by nothing. henad-compute::gpu::primitives::reduce's leaf convention is copy-pasted per model —gpu_boids/reduce.wgslandgpu_ants/reduce.wgslshare about 30 lines of identical workgroup-reduction boilerplate, differing only in howvalueis computed. A shared prelude string, or the trait supplying just the value expression, would remove the copy.
Manual notes (human)¶
The GPU ants model depends on the CPU implementation, so I'm not sure if it's a good idea. We'll probably change this when GPU model authoring is sufficiently mature.
Noticed that a lot of agents accumulate in the top-left corner of the world, which might be worth investigating. This seems to be present in the CPU implementation as well, so perhaps we did something wrong there? Since we implemented our model as a reference implementation for the krABMaga one, this might also be an issue on their part.
Update: Alright it did turn out to be their problem. I asked an agent to document this.