The reboot-survival gate is a test no operator can stand inside: the machine under test is the machine running the session. So the machine records its own recovery. openfut-boot-evidence.service polls until the anchor, Core and host agree on a namespace (or a 180s deadline expires), then writes a JSON file with the boot id, anchor/Core/host pids and netns inodes, unit states, restart counts, whether the reconciler had to act this boot, mount count, four non-mutating probes, route ownership, and the full economy snapshot — plus the journal for the boot so ordering is read from real timestamps rather than inferred from unit dependencies. It observes only; it never starts, stops or repairs anything, and carries no Requires= or ordering that anything else waits on, so it cannot affect the boot it is measuring. If the chain is broken the file says so, which is the point. Polling rather than a fixed sleep means a boot-time reconcile retry is recorded as "settled late" rather than as a failure.
OpenFUT systemd units
Supervision for OpenFUT Core and the FIFA17 UTAS host. Replaces the previous
setsid nohup … nsenter … launch, which had no restart policy, no boot
persistence and no supervisor-visible logs.
| file | scope | installed? |
|---|---|---|
openfut-staging-netns.service |
binds the staging anchor netns | yes — proving ground |
openfut-staging-core.service |
staging Core, port 18081 | yes |
openfut-staging-host.service |
staging host, port 8299 | yes |
openfut-staging-netns-reconcile.{service,timer} |
staging netns lifecycle | yes |
openfut-netns.service |
binds the production anchor netns to /run/netns/openfut |
template only |
openfut-core.service |
production Core, port 18080 | template only |
openfut-host.service |
production host, port 8099 | template only |
openfut-netns-reconcile.{service,timer} |
production netns lifecycle | template only |
openfut-netns-bind.sh |
resolves the anchor netns by NAME, idempotently, drains stale mounts | helper |
openfut-netns-reconcile.sh |
keeps services in the anchor's CURRENT netns | helper |
openfut-wait-tcp.sh |
bounded readiness gate | helper |
openfut-supervision-install.sh |
install / start / status / uninstall per environment | helper |
Production templates are not installed. Deploy only via the plan in
OpenFUT-Vault/06 Operations/OpenFUT Service Supervision (staging-proven).md.
The three decisions worth knowing
1. Wants=, not Requires=, from host → Core. Measured on staging:
Requires propagates a Core stop into a host stop, and a later Core start does
not bring the host back — a routine Core restart would leave the client with
no server. With Wants, the host survives a Core outage, answers
503 core_unavailable (never a Python fallback), and resumes the moment Core
returns, with no supervisor intervention.
2. Readiness is an ExecStartPre TCP gate, not ordering. After=/Wants=
order units; Type=exec only proves the binary exec'd. Neither means Core can
serve. Core binds its listener after opening the DB, running migrations and
loading the content pack, so "port open" is a genuine readiness signal. The gate
is bounded and fails rather than blocking: a host that waits forever looks
healthy to the supervisor while serving nobody.
3. The netns is resolved by container NAME at every start. Production must
run inside openfut-fut-backend's network namespace. The container is
restart=unless-stopped, and its netns inode changes on restart — observed
net:[4026539938] → net:[4026540033]. A hardcoded pid is therefore wrong by
construction, and any process left in the old namespace keeps running with no
interfaces, silently serving nobody. openfut-netns-bind.sh re-resolves by name
and refreshes a stale bind mount; NetworkNamespacePath= then enters it
declaratively.
4. Stale namespaces are repaired automatically (this closes the old gap).
Recreating the anchor strands already-running services in the dead namespace,
and — the dangerous part — systemd still reports them active. Measured before
the fix: anchor net:[4026539938] → net:[4026540033], Core and host unchanged in
the old one, both units active, traffic ConnectionResetError.
openfut-netns-reconcile.sh on a 10s timer compares the namespace the services
are actually in against the anchor's current one and, only on a real
difference, performs one stop → rebind → start cycle. It is level-triggered, so
it cannot miss an event and needs no debounce: a burst of three back-to-back
recreations produced exactly one rebind.
Operating
sudo ./openfut-supervision-install.sh status staging # units, anchor, netns agreement, mounts
journalctl -u openfut-staging-netns-reconcile -f # silent unless it acts
sudo systemctl restart openfut-staging-core # host survives and recovers
sudo ./openfut-supervision-install.sh install staging # copy + enable (does not start)
sudo ./openfut-supervision-install.sh uninstall staging # disable, stop, remove units
status is the one command worth knowing: it prints each unit's state, the
anchor's current netns, the namespace each service is actually in with an
ok/MISMATCH verdict, and the mount count (1 healthy, >1 a leaked stack).
Config lives in EnvironmentFiles (…/systemd/core.env, host.env), generated
from the live process environment so supervision changed how the processes
start and nothing about what they do. Binaries are immutable copies, so a
later cargo build cannot change what is running.
Interaction with the staging lifecycle script
scripts/sold-staging-up.py still starts its own unsupervised processes and
refuses to run while a staging stack is up. Stop the units first:
sudo systemctl stop openfut-staging-host openfut-staging-core
python3 scripts/sold-staging-up.py --club real --variant highest …
Reconciling the two (having the script drive the units) is deliberately out of scope for the supervision milestone.