Files
OpenFUT/scripts/systemd/openfut-netns.service
T
funman300 8ae432223a ops: systemd supervision for Core and the FIFA17 host (staging-proven)
Replaces the detached `setsid nohup … nsenter …` launch, which had no restart
policy, no boot persistence and no supervisor-visible logs. Staging units are
installed and proven; production units are TEMPLATES and are not installed.

Three decisions, each measured rather than assumed:

* `Wants=`, not `Requires=`, from host to Core. With `Requires`, stopping Core
  stopped the host AND a later Core start did not bring it back -- a routine
  Core restart would leave the client with no server. With `Wants` the host
  survives a Core outage, answers 503 core_unavailable, never falls back to
  Python, and resumes the moment Core returns with no intervention. Both halves
  tested.
* Readiness is a bounded ExecStartPre TCP gate, because ordering proves nothing
  about readiness and Type=exec only proves the binary exec'd. Core binds its
  listener after migrations and content load, so "port open" is a real signal.
  The gate FAILS rather than blocking: a host that waits forever looks healthy
  while serving nobody.
* The netns is resolved by container NAME every start. The container is
  restart=unless-stopped and its netns inode CHANGES on restart (measured:
  4026539938 -> 4026540033), so a hardcoded pid is wrong by construction and
  anything left in the old namespace serves nobody. Proven equivalent to today's
  nsenter against a scratch container, never production's namespace.

`systemd-analyze verify` caught two real defects before deployment:
StartLimitIntervalSec/StartLimitBurst sat in [Service], where systemd 252
silently ignores them, so the crash-loop ceiling was not taking effect; and a
Documentation URL containing %20 parsed as a specifier. Both fixed and the
effective properties re-confirmed from the running units.

Staging evidence: Core-first ordering, host refused when Core is absent or
merely not listening, outage survival, automatic recovery, restart, graceful
stop with no strays, boot simulated via multi-user.target, 3x SIGKILL contained
at ~5s spacing, journald logs, and economy state byte-identical throughout
(integrity ok, fk 0).
2026-08-22 20:46:15 +00:00

30 lines
1.1 KiB
Desktop File

[Unit]
Description=OpenFUT: publish the FIFA17 container network namespace to /run/netns
Documentation=file:///home/alex/OpenFUT/scripts/systemd/openfut-netns-bind.sh
# PRODUCTION TEMPLATE — NOT INSTALLED. Staging runs in the host netns and needs
# none of this; only production enters the `openfut-fut-backend` container's
# namespace.
After=docker.service
Requires=docker.service
[Service]
Type=oneshot
RemainAfterExit=yes
# Re-resolves the container by NAME on every start, so no pid is ever hardcoded
# and a container restart is repaired by restarting this unit. It exits non-zero
# when the container is not running, which is what makes Core's Requires= on it
# a real admission gate rather than decoration.
ExecStart=/home/alex/OpenFUT/scripts/systemd/openfut-netns-bind.sh openfut-fut-backend openfut
# Deliberately NO ExecStop unmount. Tearing the bind mount down while Core and
# the host are still inside that namespace would strand them; the namespace is
# owned by the container's lifetime, not by this unit's.
StandardOutput=journal
StandardError=journal
SyslogIdentifier=openfut-netns
[Install]
WantedBy=multi-user.target