Implementation progress: `specs/qemu-test-vms.md`
On Fri, 25 Sep 2026, by @lucasdicioccio, 2905 words, 3 code snippets, 0 links, 0images.
Generated from specs/qemu-test-vms-progress.md — the repository is the canonical source, and may be ahead of this page.
Implementation progress: specs/qemu-test-vms.md
Status: living doc, update as work continues. Companion to
specs/qemu-test-vms.md (the design) — this file tracks what’s actually
built, what’s still untested, and exactly what to do next. Read the spec
first if you need the “why.”
Work on this branch is being committed incrementally as it lands (see
git log); this doc may still lag the latest working-tree state by a
change or two at any given moment.
0.3 Headline: fixed the flake — two real teardown bugs, both pre-existing (2026-09-08)
Following a flake reported in §0.2 (Postgres replication VM test timing out
only inside the full 20-test suite, not standalone): the actual cause was
found by inspecting the live machine rather than guessing from logs — ps
showed 6 orphaned qemu-system-x86_64 processes left running from
earlier interrupted/crashed test runs, eating ~3GB RAM and real CPU, which
starved the timing-sensitive replication test under full-suite load.
Cleaning them up and rerunning made the flake disappear, but the real fix
was finding why teardown wasn’t cleaning them up — two separate,
pre-existing bugs, neither introduced by §0.2’s privilege work:
LinuxBridge.tapdeclared the shared bridge as its own graph dependency (deps [bridge ...]).downTree’s “release a predecessor once its last dependent is torn down” rule is correct in general (the directory/two-files case documented in CLAUDE.md) but wrong here: the bridge is explicitly meant to persist across many taps/VMs coming and going, and each tap’sdownTreecall has no visibility into other taps still relying on the same bridge (different process, different traversal). Result: tearing down any single VM deleted the shared bridge out from under every other still-running VM, leaving their tapsNO-CARRIER— directly observed on the live machine (ip link show salmontest0→ “Device does not exist” while several taps sat orphaned). Fixed by dropping the graph dependency entirely and instead runningbridge’s ownOpvia a nestedupTreeinsidetap’sup(same accepted “nested traversal, check theBool,throwIOifFalse” pattern asPostgresMigrations.remoteMigrateOpaqueSetup—howto-ops.md§5) — brings the bridge up as a precondition without ever making it this tap’s own teardown-reachable predecessor.Systemd.systemdServicenever actually set adownaction at all — it defaulted toExtension’s no-op, silently contradictingQemu.hs’s own module haddock, which already claimed “down goes through plainsystemctl stop”. This is why the orphaned VMs existed in the first place: every singlewithVmAtteardown, even a clean, successful one, left the qemu process running forever —downdeleted the tap and the unit file (viaconfigContents’s own realdown) but never stopped the service itself. Fixed by adding aStopSystemCtlCalland wiringdown = stop— confirmed viaps/systemctl --user list-unitsimmediately after a test run: no leftover process, no leftover unit file, where before there reliably was one of each perwithVmAtcall.
Confirmed via ps/ip link/systemctl --user list-units before and
after: the full 20-test suite (cabal test salmon-ops-recipes) passed
clean twice in a row post-fix (168s, 186s), leaving zero qemu processes
and zero unit files behind either time — no flake recurrence.
0.2 Headline: the whole tier runs unprivileged now, no sudo at all (2026-09-01/08)
Following up on §4’s privilege open question: dropped the requirement that
the whole test binary run as root. Test.Harness.hasVmPrivileges now
accepts either real root or a one-time capability grant; withVmAt runs
qemu as the invoking user (via LinuxBridge.Tap’s tapOwner and
Qemu.VmConfig’s vm_user/vm_group) instead of root:root. New
Salmon.Builtin.Nodes.Capabilities (setcap/getcap, idempotent) plus a
salmon-qemu-host-setup-fixture executable do the one-time host grant as
a real Op graph instead of a shell snippet — see its own haddock for the
exact commands. Confirmed passing fully unprivileged: QemuSmokeSpec
(20s) and PostgresReplicationSpec (69s, two VMs) both green with no
sudo anywhere in the test invocation.
Two real, boot-validated bugs found getting there, neither guessable from code review:
- Granting
cap_net_admintoipitself does not work.straceon a failing unprivilegedip link add ... type bridgeshowedipunconditionally callingcapset({...}, {effective=0, permitted=0, inheritable=0})at startup — iproute2 drops its entire capability set on exec and only trusts the ambient set afterwards, which a plain file-capability grant can never populate (the kernel zeroes ambient for any exec of a “privileged” file, by design). Confirmed via a clean control test first:ping(file-capcap_net_raw) works fine unprivileged, proving the capability mechanism itself was never the problem. Fix: grant the capability tocapshinstead, and haveipinvocations go throughcapsh --inh=cap_net_admin --addamb=cap_net_admin -- -c "ip ...args..."(Salmon.Builtin.Nodes.LinuxBridge.ipLinkCommand) — ambient capabilities do propagate across exec and are what iproute2 actually honors. Raising ambient itself needs the capability in both the process’s permitted and inheritable sets (--inh=first) since exec does not carry a file’s inheritable bit into the new process’s own inheritable set. Works identically for real root (whose permitted set is already full) and for an unprivileged user with the grant oncapsh, so the wrapping is unconditional now, not privilege-mode -specific. - A qemu VM’s systemd unit can’t live under
/etc/systemd/systemunprivileged (plain permission denied writing there). Fix: addedSystemd.Scope(System/User) toSalmon.Builtin.Nodes.Systemd;Qemu.VmConfiggainedvm_systemd_scope/vm_unit_dir, andTest.Harness.withVmAtnow usesSystemd.Useragainst a resolved~/.config/systemd/user, withsystemctl --useranddefault.targetswapped in formulti-user.target/network-online.target(which don’t exist in the user manager). Systemd also rejectsUser=/Group=in a user-manager unit, sorender_serviceomits them forUserscope.PgBouncer/Postgrest/MicroDNS’s existingSystemd.Configcall sites were updated to explicitSystem//etc/systemd/system— no behavior change for them.
Not yet chased down: the Postgres replication VM test flaked once when run
as part of the full 20-test suite (cabal test salmon-ops-recipes) —
walsender process due to replication timeout inside the guest, a
stale-pidfile postgres restart loop — while passing cleanly twice
standalone. Smells like timing/resource contention from running
back-to-back with everything else rather than a regression from the
privilege changes above (this project already has one documented
unrelated Layer 2 podman flake under load), but not confirmed either way.
0.1 Headline: Phase 5’s real recipe test now passes for real (2026-08-21)
Test.PostgresReplicationSpec (see §1) now passes under sudo against real
pg-primary/pg-standby rootfses — closes §3 item 4. Getting from “builds
clean, never run” to a real pass took four more bugs, none guessable without
actually booting the VMs:
resolveFixtureBinarytook the first line ofcabal list-bin’s stdout, not the last: undersudo(root, no prior cabal config),cabalprints a one-line notice (“Config file path source is default config file.”) to stdout before the actual bin path, so the scp target became that notice string instead of a path. Fixed to take the last non-blank line.postgresql-clientisn’t pulled in by thepostgresqlmeta-package the waypostgresql-client-17is — the fixture’sDebian.psql/Debian.pg_ctlmap to the genericpostgresql-clientpackage name, which wasn’t in the rootfs’s--includelist, so the guest (no network route past boot, by design) failed trying to fetch it. Fixed by chroot-installing it into the master rootfs (host has network) before copying out — see §2.- The cluster ended up on port 5433, not 5432:
pg_createcluster(during thepostgresqlpackage’s postinst, run inside achrooton the host) auto-picks the next free port by probing the host’s own running processes — since it shares the host’s network stack, it saw something already on 5432 and picked 5433. The fixture hardcodesprimaryPort = 5432. Fixed by forcingport = 5432in the rootfs’spostgresql.confdirectly. Test.Harness.sshToVmsilently mis-delivers any argument containing embedded whitespace.sshjoins every argument after the destination with a single space and ships the result as one string for the remote shell to tokenize — exactly like typing the words by hand at a terminal. So anargselement that’s a whole SQL statement or acmd 2>&1redirection doesn’t arrive as one remote token: the remote shell re-splits it on spaces along with everything else. E.g.["psql", "-tAc", "SELECT state FROM pg_stat_replication;"]arrived remotely aspsql -tAc SELECT state FROM pg_stat_replication;—-tAconly capturedSELECT, andwaitForStreaming’s query silently malformed on every single poll, meaning it could never have detected real streaming state regardless of whether replication actually worked. Same class of bug as theSystemd.render_startExecStart=quoting fix in §0 item 4 below — just in the test harness instead of production code. Fixed by addingTest.Harness.quoteForRemoteShell(single-quotes a string so it survives ssh’s space-join as one token) and applying it at each call site inTest.PostgresReplicationSpecthat needs one. Deliberately not made automatic insidesshToVmitself:Test.QemuSmokeSpec’s existingsshToVm access ["echo smoke-ok"]relies on the remote shell’s own re-splitting to turn one Haskell string into two remote words — quoting every argument unconditionally would instead hand the remote shell one literal token"echo smoke-ok"(a program name with a space in it) and break that passing test.
Confirmed passing standalone (--pattern 'Postgres replication', ~100s) and
with the rest of the non-root suite around it unaffected (Podman’s Layer 2
container test failed in one non-root sanity run, but that’s an unrelated,
pre-existing environmental flake — Salmon.Builtin.Nodes.Podman wasn’t
touched by any of this work).
0. Headline: the Layer 3 tier now works end to end (2026-08-20)
Test.QemuSmokeSpec (new) boots a real qemu VM via Test.Harness.withVm
against a hand-built debootstrap rootfs and SSHes into it for real —
confirmed passing under sudo (root is required, see §2):
cabal build salmon-ops salmon-ops-recipes:test:salmon-ops-recipes-test
sudo PATH="$PATH" dist-newstyle/build/x86_64-linux/ghc-9.8.2/salmon-ops-recipes-0.1.0.0/t/salmon-ops-recipes-test/build/salmon-ops-recipes-test/salmon-ops-recipes-test --pattern Qemu
# Qemu (Layer 3, real VM boot via withVm)
# boots the smoke rootfs and answers SSH: OK
Getting there took four real production bugs, each hand-validated by booting an actual VM and reading its serial console — none of these were guessable from code review alone:
- NIC naming (predicted in §3.1 below, confirmed for real): the
virtio-net device came up as
ens4, noteth0, breaking theip=kernel arg silently. Fixed by addingnet.ifnames=0 biosdevname=0toQemu.kernelCmdlineunconditionally (this whole tier already assumes a single, always-eth0NIC). - 9p root never mounts: the stock debootstrap initrd never even
attempts a 9p mount of its own root (
9pnet/9pnet_virtio/9pare kernel modules, not builtin, and nothing loads them) — panics with/dev/root does not exist. Fixed with a new op,Debootstrap.ensureVm9pBoot, that appends those modules to/etc/initramfs-tools/modulesand regenerates the initrd via a chroot. Also required renaming the 9p mount tag from/dev/roottovroot(androot=vrooton the cmdline):initramfs-tools’slocal_device_setuponly skips its udev block-device wait for aROOTthat neither starts with/devnor contains=— anything else, 9p tags included, it waits forever since 9p never produces a udev block device. security_model=mappedbreaks/sbin/init:run-init: /sbin/init: Too many symbolic links encountered—mappeddoesn’t round-trip Debian’s/bin -> usr/bin-style symlinks faithfully. Since qemu (and this whole tier) already runs as root on the host, switched tosecurity_model=passthrough(real symlinks/ownership preserved, no uid remapping) inQemu.qemuArgs.Systemd.render_startnever quotesExecStart=args (found only once #1–#3 above were fixed and the VM booted but SSH still never answered): it joins args with a bareText.unwords.-append’s value is one multi-word string; unquoted in the unit file, systemd’s ownExecStart=parser splits it back into several separate qemu arguments, so the kernel cmdline silently never arrives intact. Fixed by quoting any arg containing whitespace/shell metacharacters inSystemd.hs. This is a generalSystemd.hsbug, not qemu-specific — just never exercised before since nothing else here passes a multi-word singleExecStart=argument.
A fifth issue was in the test setup, not production code: the original
plan (§4 step 2, now superseded) had a human manually copying their own
~/.ssh/id_ed25519.pub into the rootfs’s authorized_keys. Running the
privileged tier under sudo doesn’t forward the invoking user’s
ssh-agent, so pubkey auth via a personal key silently never succeeds and
waitForSsh just times out. Fixed by having withVm generate its own
ephemeral SSH CA + signed client key per boot (Test.Harness. ensureVmSshAccess, using the existing but previously-unused
Keys.sshKey/Keys.signKey CA-signing primitives) and provision the
guest’s sshd to trust it (TrustedUserCAKeys + PasswordAuthentication no drop-in) — no manual key step needed any more. This surfaced one more
bug along the way: Keys.signKey never passed -n <principal> to
ssh-keygen -s, and modern OpenSSH (checked against 9.6p1) hard-rejects a
certificate with an empty principal list (Certificate lacks principal list) — contrary to older folklore that an empty list means “valid for
any principal.” Fixed by adding a [Principal] parameter to signKey
(safe: grep confirmed nothing else in the codebase called it yet).
1. Files touched so far
| File | State | What |
|---|---|---|
salmon-ops/src/Salmon/Builtin/Nodes/LinuxBridge.hs | new | bridge, tap, bridgeAddr ops. |
salmon-ops/src/Salmon/Builtin/Nodes/Qemu.hs | new | VmConfig, resolveKernelInitrd, qemuArgs, setup. Boot-validated fixes: security_model=passthrough, mount_tag=vroot/root=vroot, net.ifnames=0 biosdevname=0 baked into kernelCmdline, -cpu host paired with -enable-kvm. |
salmon-ops/src/Salmon/Builtin/Nodes/Debian/Debootstrap.hs | modified | vmEssentials :: Includes (kernel + openssh-server), plus new ensureVm9pBoot op (9p initramfs modules + update-initramfs via chroot). |
salmon-ops/src/Salmon/Builtin/Nodes/Systemd.hs | modified | render_start now quotes ExecStart= args containing whitespace/shell metacharacters — real bug fix, not qemu-specific. |
salmon-ops/src/Salmon/Builtin/Nodes/Keys.hs | modified | signKey gained a required [Principal] parameter (-n to ssh-keygen -s) — modern OpenSSH rejects principal-less certs. |
salmon-ops/salmon-ops.cabal | modified | registered LinuxBridge and Qemu in exposed-modules. |
salmon-ops-recipes/test/Test/Harness.hs | modified | Layer 3 section: testBridge/testBridgeCidr/testVmAddr/testVmAddr2/ensureTestBridge/withVm/withVmAt, VmAccess/sshToVm/scpToVm (replaces the old bare-Ssh.Remote interface — see §0 item 5), ensureVmSshAccess. withVm is now withVmAt testVmAddr; withVmAt takes the guest address as a parameter so more than one VM can be up at once on the shared test bridge (nested calls, one address each) — added for §3 item 4. quoteForRemoteShell added 2026-08-21 (see §0.1 item 4) — single-quotes a caller's sshToVm argument so it survives ssh's own space-join as one remote token; not applied automatically inside sshToVm itself (would break Test.QemuSmokeSpec's existing ["echo smoke-ok"] call, which relies on the old join-then-resplit behavior). |
salmon-ops-recipes/test/Test/QemuSmokeSpec.hs | new | Layer 3 smoke test: boots the smoke rootfs via withVm, asserts SSH answers. Skips loudly (not fail) without root, without qemu-system-x86_64, or without a pre-built rootfs at /var/lib/salmon-test-vms/smoke/root. |
salmon-ops-recipes/test/Test/DebootstrapSpec.hs | new | Layer 3: runs Debootstrap.rootTree inject Debootstrap.ensureVm9pBoot as real Ops (not the equivalent hand-run chroot script) against /var/lib/salmon-test-vms/debootstrap-op-smoke/root, checks the 9p modules got written, then reruns once more to confirm idempotency. Closes §3 item 2. Gated behind root/debootstrap-on-PATH (skip loudly) plus, for the first (network-heavy) run only, the opt-in env var SALMON_TEST_RUN_DEBOOTSTRAP=1 — deliberately does not wipe the rootfs between runs, so once debootstrapped once, every later run is offline/seconds-long via the two ops' own prelims. Confirmed passing under sudo (1562.57s first run, network-bound). |
salmon-ops-recipes/test/Test/PostgresReplicationSpec.hs | new | Layer 3 port of the hand-run salmon-ops/fixtures/PostgresReplicationFixture.hs: boots a primary VM (testVmAddr) and a standby VM (testVmAddr2) via withVmAt, scpToVms the already-built fixture binary (resolved via cabal list-bin, no hardcoded path) onto each, drives it over SSH exactly like the old fixture's manual podman exec steps, then — unlike the old fixture, which just told a human to eyeball psql — polls pg_stat_replication for streaming, inserts a row on the primary, and polls the standby until that row actually shows up. Closes §3 item 4 (2026-08-21) — confirmed passing under sudo for real (~100s), after the four bugs in §0.1. On a failed fixture run or a waitForStreaming timeout, dumps pg_lsclusters/postgres logs/pg_stat_replication/pg_stat_wal_receiver/a standby→primary ping into the failure message (failure-path only, no extra SSH round-trips on the success path). |
salmon-ops-recipes/test/Test/QemuResolveKernelSpec.hs | new | Layer 1, no root/VM needed: builds a scratch boot/ dir via temporary's withSystemTempDirectory and checks Qemu.resolveKernelInitrd's three cases (one match resolves, zero throws, two — a held-over old kernel — throws "ambiguous" rather than silently picking one). Closes §3 item 3 (2026-08-21). |
salmon-ops-recipes/test/Main.hs, salmon-ops-recipes.cabal | modified | wired Test.QemuSmokeSpec, Test.DebootstrapSpec, Test.PostgresReplicationSpec, Test.QemuResolveKernelSpec in; added unix to test-suite build-depends (for getEffectiveUserID). |
cabal build salmon-ops salmon-ops-recipes salmon-ops-recipes:test:salmon-ops-recipes-test
and cabal test salmon-ops-recipes (non-root; all Layer-3-needing tests
vacuously skip, everything else including a real Layer 2 podman test
passes — 17 tests, all green) both clean as of 2026-08-20. The sudo-run
Qemu and Debootstrap tests are the real, non-vacuous confirmations (see
§0, §3 item 2); the Postgres replication test is written and building but
not yet run for real (see its row above).
2. Environment state (this machine, checked 2026-08-20)
iproute2(ip),qemu-system-x86(providesqemu-system-x86_64),debootstrap(1.0.134ubuntu2): all installed.- KVM available and now exercised:
/dev/kvmexists, 40vmx/svmflags in/proc/cpuinfo.Test.Harness.withVmnow defaults tovm_enable_kvm = Truewith-cpu host(see §3.1 for the fix history); confirmed passing undersudo, boot+SSH completing in single-digit seconds in the two runs measured so far, vs. ~80–140s for the earlier TCG-only runs. - The smoke rootfs lives at
/var/lib/salmon-test-vms/smoke/root, built via (see §0 item 2 for whyensureVm9pBoot’s initramfs fix also needs applying — the rootfs on disk currently has that fix hand-applied via chroot, not yet re-derived by actually running the newDebootstrap.ensureVm9pBootop against it):
Nosudo debootstrap --include=linux-image-amd64,openssh-server stable /var/lib/salmon-test-vms/smoke/rootauthorized_keysprovisioning needed any more (§0 item 5) —withVmhandles its own access. - A second rootfs,
/var/lib/salmon-test-vms/debootstrap-op-smoke/root, built byDebootstrap.rootTree/Debootstrap.ensureVm9pBootthemselves (viaTest.DebootstrapSpec, see §1) rather than by hand — this is the one that actually proves those ops work, as opposed to the smoke rootfs above which still carries a hand-applied 9p fix. /var/lib/salmon-test-vms/pg-primary/rootand/var/lib/salmon-test-vms/pg-standby/root(§1) built 2026-08-21 — postgres is baked in at debootstrap time rather than apt-installed inside the guest at test time, since the test bridge has no NAT/internet route out of a guest past boot. Rather than debootstrapping each separately (they need the identical package set), built once as apg-masterrootfs and copied out twice:
Safe here specifically becausesudo debootstrap --include=linux-image-amd64,openssh-server,postgresql,sudo stable /var/lib/salmon-test-vms/pg-master/root # apply Debootstrap.ensureVm9pBoot's fix (see §0 item 2) to pg-master once # chroot-install postgresql-client into pg-master (see §0.1 item 2 — not # pulled in by the `postgresql` meta-package the way postgresql-client-17 is) # force `port = 5432` in pg-master's postgresql.conf (see §0.1 item 3 — # pg_createcluster auto-picked 5433 since the chroot install shares the # host's own network stack/port-in-use probing) sudo rsync -aHAX --numeric-ids /var/lib/salmon-test-vms/pg-master/root/ /var/lib/salmon-test-vms/pg-primary/root/ sudo rsync -aHAX --numeric-ids /var/lib/salmon-test-vms/pg-master/root/ /var/lib/salmon-test-vms/pg-standby/root/Test.Harness.sshToVmusesStrictHostKeyChecking=no/UserKnownHostsFile=/dev/null(no host-key verification, so duplicate host keys across the two copies don’t matter), andpostgresql’s postinst doesn’t start the service or write instance-specific state during a chroot debootstrap (services don’t autostart in a chroot) — the copied data directories are just the vanilla package-created default cluster; primary vs. standby role is entirely a runtime distinction made by the fixture, not baked into the rootfs.
3. What’s still open
- ~~KVM never actually exercised end to end.~~ Done (2026-08-20).
Flipped
Test.Harness.withVm’sQemu.vm_enable_kvmtoTrueand reranQemuSmokeSpecundersudo— passed. First KVM pass was noticeably slower than the proven TCG baseline; root cause:-enable- kvmwas passed without-cpu host, so the guest still ran the generic emulatedqemu64CPU model — KVM only pays off once the guest actually gets a KVM-aware CPU model. Fixed inQemu.qemuArgs:if cfg.vm_enable_kvm then ["-enable-kvm", "-cpu", "host"] else []. Rerun after the fix passed in ~5.5s (one earlier run hit ~3.75s) — both far faster than the ~80–140s TCG figure recorded in §2, though with only two data points this could partly be host-side caching rather than a clean KVM-vs-TCG comparison; not worth chasing further unless boot time becomes a problem again.vm_enable_kvm = Trueis now the harness default. - ~~
Debootstrap.ensureVm9pBootitself is untested as a salmonOp.~~ Done (2026-08-20).Test.DebootstrapSpecrunsDebootstrap.rootTreeinjectDebootstrap.ensureVm9pBootfor real against a fresh rootfs, checks the 9p modules land, and reruns once more to confirm both ops’prelims reportSkippablethe second time — passed undersudo(1562.57s, almost entirelydebootstrap’s own package downloads). Gated behindSALMON_TEST_RUN_DEBOOTSTRAP=1for that first real run and does not wipe the rootfs between runs, so it doesn’t silently re-fetch ~100+ packages (and burn a metered connection) on every suite run — see its row in §1. - ~~
resolveKernelInitrd’s prefix match~~ Done (2026-08-21). AddedTest.QemuResolveKernelSpec— a Layer 1 (pure filesystem, no root/VM/debootstrap) test that builds a scratchboot/dir viawithSystemTempDirectoryand checks all three cases: exactly-one-match resolves, zero matches throws"no vmlinuz-*", and two matches (simulating a held-over old kernel) throws"ambiguous vmlinuz-*"rather than silently picking one. All three pass. No rootfs/VM bugs found —resolveKernelInitrd’s existing ambiguity handling was already correct, just previously unverified. - ~~Phase 5 of the spec: pick a real recipe Layer 2 can’t exercise well and
write its first real Layer 3 test.~~ Done (2026-08-21).
Test.PostgresReplicationSpec(§1) portssalmon-ops/fixtures/PostgresReplicationFixture.hs’s hand-run, eyeballed podman flow onto two real qemu VMs with actual pass/fail assertions (pg_stat_replicationreachesstreaming, a row written on the primary shows up on the standby) — confirmed passing undersudofor real (~100s), after the four bugs in §0.1 (acabal list-binoutput-parsing bug, a missingpostgresql-clientpackage, a wrong postgres port, and a silentsshToVmargument-quoting bug affecting every multi-word remote command this test ran). Needed generalizingwithVmintowithVmAt(a caller-chosen guest address) so two VMs can be up at once on the shared test bridge, plus a newscpToVmto get the compiled fixture binary onto each guest.
4. Design decisions made while implementing (not equally emphasized in the spec)
bridgeAddr/CidrinLinuxBridge.hs— not in the original spec write-up, added because the host side of the test bridge needs an address for SSH to route through. Same idempotency shape asbridge/tap(ip addr show devgrep viaprelim).- Test subnet:
10.99.0.0/24, host (bridge)10.99.0.1, guest fixed at10.99.0.2(Test.Harness.testBridgeCidr/testVmAddr) — single-VM- at-a-time assumption for v1, matching the spec’s “prove the tier end to end” framing before anything like an address pool. - Bridge name:
salmontest0, left standing across test runs (persistent, matching the spec’s leaning in its bridge-lifecycle open question).
- Test subnet:
Qemu.setup’s systemd unit runs qemu asroot:root(Test.Harness.withVmsetsvm_user/vm_group = "root") rather than threading throughLinuxBridge.tap’stapOwnerfor an unprivileged user — simplest given this whole tier already assumes privileged execution (documented prerequisite, per the spec’s privilege open question), avoids a second permissions mechanism to get right on the first pass. This also motivated thesecurity_model=passthroughchoice in §0 item 3.- Graceful VM shutdown via the qemu monitor socket (spec §2) is not
implemented —
downgoes through plainsystemctl stop, i.e. SIGTERM. Explicitly called out as acceptable for v1’s disposable-VM use case inQemu.hs’s module haddock; revisit if abrupt termination ever causes a real problem (e.g. corrupting guest filesystem state between runs). - Test SSH access is now entirely
withVm’s own responsibility (a fresh per-boot CA + signed key, see §0 item 5) rather than something the caller pre-provisions — a deliberate narrowing from the original “caller suppliesauthorized_keys” plan once the sudo/agent problem showed up in practice. Production recipes (e.g. a real CA-backed service) still follow the project’s key-exchange-agnostic convention; this only changes how the test harness itself authenticates to its own disposable, harness-owned VM.