Files
OpenFUT/scripts/systemd
funman300 4156dc5810 ops(systemd): record unattended proof of post-boot recovery
The reboot-survival gate is a test no operator can stand inside: the machine
under test is the machine running the session. So the machine records its own
recovery.

openfut-boot-evidence.service polls until the anchor, Core and host agree on
a namespace (or a 180s deadline expires), then writes a JSON file with the
boot id, anchor/Core/host pids and netns inodes, unit states, restart counts,
whether the reconciler had to act this boot, mount count, four non-mutating
probes, route ownership, and the full economy snapshot — plus the journal for
the boot so ordering is read from real timestamps rather than inferred from
unit dependencies.

It observes only; it never starts, stops or repairs anything, and carries no
Requires= or ordering that anything else waits on, so it cannot affect the
boot it is measuring. If the chain is broken the file says so, which is the
point.

Polling rather than a fixed sleep means a boot-time reconcile retry is
recorded as "settled late" rather than as a failure.
2026-08-22 21:38:28 +00:00
..

OpenFUT systemd units

Supervision for OpenFUT Core and the FIFA17 UTAS host. Replaces the previous setsid nohup … nsenter … launch, which had no restart policy, no boot persistence and no supervisor-visible logs.

file scope installed?
openfut-staging-netns.service binds the staging anchor netns yes — proving ground
openfut-staging-core.service staging Core, port 18081 yes
openfut-staging-host.service staging host, port 8299 yes
openfut-staging-netns-reconcile.{service,timer} staging netns lifecycle yes
openfut-netns.service binds the production anchor netns to /run/netns/openfut template only
openfut-core.service production Core, port 18080 template only
openfut-host.service production host, port 8099 template only
openfut-netns-reconcile.{service,timer} production netns lifecycle template only
openfut-netns-bind.sh resolves the anchor netns by NAME, idempotently, drains stale mounts helper
openfut-netns-reconcile.sh keeps services in the anchor's CURRENT netns helper
openfut-wait-tcp.sh bounded readiness gate helper
openfut-supervision-install.sh install / start / status / uninstall per environment helper

Production templates are not installed. Deploy only via the plan in OpenFUT-Vault/06 Operations/OpenFUT Service Supervision (staging-proven).md.

The three decisions worth knowing

1. Wants=, not Requires=, from host → Core. Measured on staging: Requires propagates a Core stop into a host stop, and a later Core start does not bring the host back — a routine Core restart would leave the client with no server. With Wants, the host survives a Core outage, answers 503 core_unavailable (never a Python fallback), and resumes the moment Core returns, with no supervisor intervention.

2. Readiness is an ExecStartPre TCP gate, not ordering. After=/Wants= order units; Type=exec only proves the binary exec'd. Neither means Core can serve. Core binds its listener after opening the DB, running migrations and loading the content pack, so "port open" is a genuine readiness signal. The gate is bounded and fails rather than blocking: a host that waits forever looks healthy to the supervisor while serving nobody.

3. The netns is resolved by container NAME at every start. Production must run inside openfut-fut-backend's network namespace. The container is restart=unless-stopped, and its netns inode changes on restart — observed net:[4026539938] → net:[4026540033]. A hardcoded pid is therefore wrong by construction, and any process left in the old namespace keeps running with no interfaces, silently serving nobody. openfut-netns-bind.sh re-resolves by name and refreshes a stale bind mount; NetworkNamespacePath= then enters it declaratively.

4. Stale namespaces are repaired automatically (this closes the old gap). Recreating the anchor strands already-running services in the dead namespace, and — the dangerous part — systemd still reports them active. Measured before the fix: anchor net:[4026539938] → net:[4026540033], Core and host unchanged in the old one, both units active, traffic ConnectionResetError.

openfut-netns-reconcile.sh on a 10s timer compares the namespace the services are actually in against the anchor's current one and, only on a real difference, performs one stop → rebind → start cycle. It is level-triggered, so it cannot miss an event and needs no debounce: a burst of three back-to-back recreations produced exactly one rebind.

Operating

sudo ./openfut-supervision-install.sh status staging    # units, anchor, netns agreement, mounts
journalctl -u openfut-staging-netns-reconcile -f        # silent unless it acts
sudo systemctl restart openfut-staging-core             # host survives and recovers
sudo ./openfut-supervision-install.sh install staging   # copy + enable (does not start)
sudo ./openfut-supervision-install.sh uninstall staging # disable, stop, remove units

status is the one command worth knowing: it prints each unit's state, the anchor's current netns, the namespace each service is actually in with an ok/MISMATCH verdict, and the mount count (1 healthy, >1 a leaked stack).

Config lives in EnvironmentFiles (…/systemd/core.env, host.env), generated from the live process environment so supervision changed how the processes start and nothing about what they do. Binaries are immutable copies, so a later cargo build cannot change what is running.

Interaction with the staging lifecycle script

scripts/sold-staging-up.py still starts its own unsupervised processes and refuses to run while a staging stack is up. Stop the units first:

sudo systemctl stop openfut-staging-host openfut-staging-core
python3 scripts/sold-staging-up.py --club real --variant highest …

Reconciling the two (having the script drive the units) is deliberately out of scope for the supervision milestone.