Skip to content

Benchmarks rerun on a rented RTX 4090

The Threadripper 3960X and RTX 4090 that produced docs/benchmarks.md is gone, so the maintainer asked for every published benchmark to be measured again on a rented RTX 4090. One secure-tier Lium pod ran all five reference engines and every Henad row, with nothing mixed in from older results. The pod needed a hand-written Vulkan driver manifest before wgpu could see the GPU, and two problems in the runbook surfaced on the way: the MASON download URL is dead, and krABMaga needs two system headers on Linux. All 24 port gates passed, the 208-run sweep finished in 6.6 hours, and Henad's own 424-configuration matrix in 6.3 more. Rankings are unchanged apart from one cell, SIR at 64², which a local A/B traced to the machine rather than to Henad's code. The pod ran for 14.2 hours and cost $6.37.

State before

master was at 3c7c7e0 with a clean tree, one commit ahead of the merge of #50 and pushed during the session. The 48-library branch named in the request no longer existed, since #50 had merged it, so 3c7c7e0 is the commit measured.

docs/benchmarks.md described a 24-core AMD Ryzen Threadripper 3960X with an RTX 4090 (host hardywen). Its figures and tables came from hardywen_20260907_perf.csv, measured at 0fd5051-dirty, and its own record (#24) asked for the sweep to be rerun. The unpublished matrix history under results/ ended at 7_bench_matrix_4090_38d0436…, from August, before the GPU network work of #46.

What was done

The pod

The candidates came from Lium's node list, filtered to secure tier and single-GPU hosts, so no other tenant shared the CPU. The node with the most RAM and disk was gentle-lion-e2 in Drobeta-Turnu Severin, Romania, at $0.45 an hour. The maintainer approved it and set a \(10 cap, so the pod ran under a 22-hour TTL (\)9.90).

CPU AMD Ryzen 9 3900, 12 cores and 24 threads, 4.36 GHz maximum, 64 MB L3, one NUMA node, bare metal
Frequency amd-pstate-epp in active mode, balance_performance, about 4.2 GHz on all 24 threads under load
Memory 94 GB, with a 90 GiB cgroup limit
Disk 937 GB NVMe, 415 GB free
GPU NVIDIA GeForce RTX 4090, 24 GB, driver 595.58.03, PCIe 4.0 x16, 450 W
Kernel and OS Linux 6.17.0-22-generic, Ubuntu 24.04.4 in Lium's default PyTorch container
Container a CPU quota of 24, and the hostname set to lium4090 so the sweep stamps a readable host

Everything lived under /workspace. The pod's /root is an encrypted FUSE mount, so HOME, CARGO_HOME, RUSTUP_HOME, the Julia depot and the uv cache all pointed there through one sourced environment file.

The repository went over as the 732 files git ls-files lists, plus .git, so the pod held exactly 3c7c7e0 and no ignored file left the maintainer's machine. rsync carried the local uid across, and git under the new HOME then refused the tree as owned by someone else. compare_bench.py would have stamped every row with commit unknown, so the tree was handed to root before anything ran.

Validating the GPU

-e NVIDIA_DRIVER_CAPABILITIES=all at creation mounted the driver's graphics libraries (libGLX_nvidia, libnvidia-glvkspirv, libnvidia-gpucomp and the rest), but no ICD manifest, and vulkaninfo found no driver. The standard manifest, written by hand and pointing at libGLX_nvidia.so.0, did not work either. The loader reported Could not get 'vkCreateInstance' via 'vk_icdGetInstanceProcAddr', and strace showed the driver probing /dev/dri, which the container lacks, without ever opening /dev/nvidiactl.

libEGL_nvidia.so.0 exports the same three ICD entry points, and a manifest pointing at it worked. vulkaninfo --summary then listed one device, the RTX 4090 on the NVIDIA proprietary driver 595.58.03 at Vulkan 1.4.329, with no llvmpipe or lavapipe. henad-cli --info reported the same adapter as a discrete GPU on the Vulkan backend. A copy of the manifest under /workspace would have restored it after a container restart. None happened.

Validating the CPU

A 75-second mpstat ran beside sysbench cpu, at one thread and at 24, twice each. Steal time was 0.00% in all 76 samples, and the idle phases showed nothing else running on the host. One thread gave 2,088 and 2,052 events a second, and 24 threads gave 26,244 and 26,231, 12.6 times one thread.

Engines

Engine Version Notes
Rust 1.97.1 from rust-toolchain
OpenJDK 21.0.12 for NetLogo and MASON
NetLogo 7.0.4 the Linux tarball, which unpacks to NetLogo 7.0.4 with a space
MASON 22 the maintainer's local jar, digest checked by fetch_mason.sh
Julia 1.12.7 through juliaup, the version Manifest.toml records
Agents.jl 7.0.3 from the committed manifest
Python 3.13.16 through uv 0.12.24
Mesa 3.5.1 from the committed lock
krABMaga 0.6.2 built by the driver, both features

Two parts of the runbook failed.

fetch_mason.sh downloaded a 55 KB HTML page and refused it on the digest. cs.gmu.edu/~eclab/projects/mason/mason.22.jar now redirects to the department's home page. The mirror at people.cs.gmu.edu serves the same jar, byte for byte against the pinned digest, so the script now fetches from there. On the pod, the maintainer's own copy of the jar was uploaded and the script verified it.

krABMaga failed to build inside the first gate run. Its plotting dependencies pull in yeslogic-fontconfig-sys and freetype-sys, which need libfontconfig1-dev and libfreetype-dev, and the prerequisites table now says so. validate_ports.py writes validated.json only at its end, so the 16 verdicts it had already earned were lost with the build, and the whole run was repeated.

Gates

All 24 gates passed on the second run, which took six minutes with the SIR replicates cached by the first.

Engine Game of Life Boids SIR Ant foraging
Mesa 3.5.1 yes yes yes yes
NetLogo 7.0.4 yes yes yes yes
MASON 22 yes yes yes yes
Agents.jl 7.0.3 yes yes yes yes
krABMaga 0.6.2 yes yes yes yes
krABMaga 0.6.2, parallel yes yes yes yes

Every fixture a passing gate writes back came out byte-identical to the committed one, apart from NetLogo's two boids fixtures, whose # generated: timestamp line changed. Their 204 numbers matched exactly. The two files were checked out again, so the sweep stamped 3c7c7e0 rather than 3c7c7e0-dirty.

The sweep

compare_bench.py --dry-run planned 208 runs, and --smoke passed 32 of 32 in under two minutes. The full sweep ran in tmux from 22:24 to 04:59 UTC, with nothing else on the pod.

Threadripper, 0fd5051-dirty Ryzen 9 3900, 3c7c7e0
ok 179 178
over budget 7 8
timeout 5 4
skipped after a slower rung 17 18
summed wall time 6.85 h 6.58 h

Mesa's boids at 3k finished five reps in 987 seconds on the Threadripper and four within the budget here, so that ladder now stops one rung earlier. Every other ladder stops on the same rung as before. NetLogo's Game of Life at 2048² managed one rep here where it managed none, which moves one row from timeout to over budget.

Henad's one-thread column sits within 17% of the Threadripper's at every rung but the three smallest. SIR runs 10 to 17% slower and boids 2 to 12% slower, as the 3900's lower boost clock predicts. Ants runs up to 15% faster, which the clock does not explain. The all-cores column has 24 threads where it had 48. Boids runs about twice as long from 3k to 300k and 1.56 times as long at 1M, and SIR 1.4 to 1.7 times as long from 1024² up. Game of Life runs faster at every rung, as do the small rungs of SIR and ants. The GPU column is the same card. Boids and SIR agree within 3% at their top rungs, and everything else runs 5 to 30% faster here, most at the small rungs, where a step costs mostly submission overhead.

The one rung that changed order

The new SIR table has Henad ahead of Agents.jl at 64², where the old one had Agents.jl ahead (0.85). Agents.jl's time there barely moved, 1.07 ms then and 1.10 ms now. Henad's fell from 1.26 ms to 0.80 ms on a slower clock, and Game of Life at 64² and 128² fell by the same factor of 1.6 to 1.8.

The code did not change at those rungs. henad-cli built at 0fd5051 and at 3c7c7e0 and run interleaved on an M4 Pro, six rounds of five reps each, came out at a new-to-old ratio of 0.95 for SIR at 64², 1.00 and 0.99 for Game of Life at 64² and 128², and 1.01 for SIR at 512². Both sweeps are steady within their own reps, with the median at most 1.2 times the minimum. So the Threadripper ran these three cells, where a whole 100-step rep takes about a millisecond, 1.6 to 1.8 times slower than this machine, and the cause is unexplained. The page has no prose about the order at that rung, and none was added.

Henad's own matrix

bench_matrix.py ran next, queued in tmux to start only if the sweep exited 0, from 04:59 to 11:15 UTC. Its output is results/8_bench_matrix_lium4090_3c7c7e0c0b4b338b36b43ac955ddd31a5f8f4077.csv, unpublished like the seven before it. The lium4090 in the name still sorts it after 7_bench_matrix_4090_… and keeps the two CPUs apart. It swept nine models, leaving out team_assembly, which the script skips unless --models names it.

August, Threadripper, 38d0436 Ryzen 9 3900, 3c7c7e0
configurations 404 424, with virus_network added
ok 392 386
timeout 12 6
error 0 32
summed wall time 7.27 h 6.27 h

The 32 errors are the 0×0 grid rung, which henad-cli now refuses (four configurations each for Game of Life and SIR, and twelve each for their GPU ports). The six timeouts are gpu_boids at a million agents and 100,000 steps. August also timed out four of the six 100k configurations over 100,000 steps, and both million-agent configurations over 1,000 steps after a 10,000-step warm-up, and all six now finish.

The GPU rows are the first RTX 4090 numbers since #46, which parked gpu_virus_network on its own branch. That model is not in the registry, so a GPU network model is still unmeasured on a 4090. The same card gives the same answer where the GPU does the work. Game of Life at 4096² and 8192², SIR from 1024² up and ants from 100k up all sit within 7% of August's throughput, at the longest and most warmed configuration of each point. Below those sizes throughput is 1.1 to 1.7 times August's, most of all at the smallest, where a step is mostly submission overhead.

The CPU rows mix a different CPU with the September changes of #24, so they say little against August. Game of Life and SIR at 64² gain 12 and 7.7 times, and ants 3 to 7.6 times, of the same order as what #24 measured for its own changes on all cores on an M4 Pro. virus_network has no earlier row to compare with. It peaks at 8.9e8 node updates a second at 100k nodes, and falls to 6.7e8 at a million.

Publishing

plot_compare.py --publish regenerated every figure and table under docs/assets/benchmarks/, the README's and the home page's headline figures among them, and copied the sweep in as lium4090_20261008.csv. Nothing in the repository names hardywen_20260907_perf.csv, so it stays where it was, beside the new CSV.

The lines-of-code table moved for Henad alone, from 57, 257, 105 and 409 to 71, 291, 130 and 446. The four model files gained their actions and Debug derives after 0fd5051, and the caveat on the page and in method.md now lists actions beside parameters, statistics and a palette.

The page names the new machine, commit and dates, and the first-model guide's "around 950 steps a second on a 24-core desktop" now reads around 1,100 on a 12-core desktop CPU, from the sweep's all-cores Game of Life at 4096². zensical build reports no issues.

crates/henad-explore/src/tests/
    replay.rs                                       a two-minute wall-clock deadline in snapshot_at, after the push
benchmarks/
    README.md                                       krABMaga's Linux headers in the prerequisites
    method.md                                       actions in the lines-of-code caveat
    mason/fetch_mason.sh                            the people.cs.gmu.edu mirror
docs/
    benchmarks.md                                   the new machine, commit and dates, actions in the caveat
    guide/first-model/game-of-life.md               the steps-a-second figure from the new sweep
    assets/benchmarks/
        *.svg                                       regenerated from the new sweep
        tables/*.snippet                            regenerated
        lium4090_20261008.csv                       new, the sweep the figures come from
    developing/agent-record/
        20261009-48-rtx-4090-rerun.md               this file
zensical.toml                                       nav entry #48

State after

The pod was removed at 11:24 UTC, after every result file was copied home and checked against the pod's copy by SHA-256, and lium ps lists no pod. It ran 14.15 hours at \(0.45, for **\)6.37** of the $10 the maintainer set aside. The sweep and the matrix took 12.9 of those hours, and setup, the gates and the checks most of the rest.

Phase UTC
Pod ready 21:15
Validated and installed 21:26
Gates, first run, stopped by the krABMaga build 21:27 to 22:13
Gates, second run, 24 of 24 22:14 to 22:20
Smoke 22:21 to 22:23
Sweep 22:24 to 04:59
Matrix 04:59 to 11:15
Pod removed 11:24

The published figures and tables come from results/compare/lium4090_20261008.csv, copied into docs/assets/benchmarks/. The gate verdicts, the smoke run, every pod log and the CPU check sit under results/compare/lium4090/, which is ignored like the rest of results/. results/compare/validated.json still holds the Threadripper's verdicts, and the pod's are in results/compare/lium4090/validated.json. The matrix is results/8_bench_matrix_lium4090_3c7c7e0….csv with its JSON beside it.

The work is three commits on master: docs: rerun benchmarks for the assets, pages and this record, docs: list krABMaga benchmark prereq for the runbook, and fix: mirror MASON jar for fetch_mason.sh. uv run --locked zensical build reports no issues, and scripts/test_docs_frontmatter.py passes.

After the push, CI's Windows job failed tests::replay::a_replayed_gpu_run_from_a_result_set_matches_its_row, a test none of the three commits touched and one that had passed on every earlier run. Its helper, snapshot_at, gave the GPU loop 1,000 polls of 10 ms to publish the target tick. The Windows runner's adapter is WARP, which renders on the same CPU as the tests running beside it, and two GPU tests in the same pass ran for over a minute each. The helper now waits until a two-minute wall-clock deadline, SNAPSHOT_DEADLINE, so a slow loop still passes and a stuck one still fails.

Issues found & future directions

  • NetLogo's fixture export stamps the time. Every NetLogo gate run rewrites two tracked files with nothing but a new timestamp, which marks the tree dirty and the sweep's commit -dirty. Dropping the line from the export, or leaving a fixture alone when only that line differs, would keep a gated tree clean.
  • One failed build discards a whole gate run. validate_ports.py builds each engine inside its loop and writes its verdicts only at the end, so a build that fails late loses every verdict before it. Writing after each engine, or reporting the build failure as that engine's verdict, would keep them.
  • Vulkan on a rented container. A container without /dev/dri gets no working Vulkan from the GLX entry point, even with every graphics library mounted. The EGL entry point works. Anyone repeating this on another host needs the manifest in the GPU section above. Without it, Vulkan finds no driver here, and on an image with Mesa's Vulkan drivers installed llvmpipe stands in, which HENAD_REQUIRE_GPU does not catch.
  • The SIR 64² cell is a property of the machine. Three cells where a rep lasts a millisecond moved by 1.6 to 1.8 times between two Zen 2 machines running the same code, and one of them changed which engine leads. A table built on rungs this small reads the machine at least as much as the engine.
  • hardywen_20260907_perf.csv is now stale. Nothing references it, and it describes figures the page no longer shows.
  • bench_matrix.py still asks for a 0×0 grid. GRID_SIZES starts at (0, 0), which ran in August, and henad-cli now refuses grid_width=0 as outside 1..=10000. Every grid model records those configurations as errors, and the 1×1 rung beside them already measures the empty-grid overhead.
  • The matrix history now spans two CPUs. plot_bench_history.py reads the host from the file name but plots every commit on one axis, so its CPU rows compare a 48-thread Threadripper with a 24-thread Ryzen as though only the commit changed.
  • The sweep takes the host from the hostname. In a container that is the container id, which is why the pod was renamed. A --host flag on compare_bench.py would record a readable name without root.

Manual notes (human)