8ae432223a
Replaces the detached `setsid nohup … nsenter …` launch, which had no restart policy, no boot persistence and no supervisor-visible logs. Staging units are installed and proven; production units are TEMPLATES and are not installed. Three decisions, each measured rather than assumed: * `Wants=`, not `Requires=`, from host to Core. With `Requires`, stopping Core stopped the host AND a later Core start did not bring it back -- a routine Core restart would leave the client with no server. With `Wants` the host survives a Core outage, answers 503 core_unavailable, never falls back to Python, and resumes the moment Core returns with no intervention. Both halves tested. * Readiness is a bounded ExecStartPre TCP gate, because ordering proves nothing about readiness and Type=exec only proves the binary exec'd. Core binds its listener after migrations and content load, so "port open" is a real signal. The gate FAILS rather than blocking: a host that waits forever looks healthy while serving nobody. * The netns is resolved by container NAME every start. The container is restart=unless-stopped and its netns inode CHANGES on restart (measured: 4026539938 -> 4026540033), so a hardcoded pid is wrong by construction and anything left in the old namespace serves nobody. Proven equivalent to today's nsenter against a scratch container, never production's namespace. `systemd-analyze verify` caught two real defects before deployment: StartLimitIntervalSec/StartLimitBurst sat in [Service], where systemd 252 silently ignores them, so the crash-loop ceiling was not taking effect; and a Documentation URL containing %20 parsed as a specifier. Both fixed and the effective properties re-confirmed from the running units. Staging evidence: Core-first ordering, host refused when Core is absent or merely not listening, outage survival, automatic recovery, restart, graceful stop with no strays, boot simulated via multi-user.target, 3x SIGKILL contained at ~5s spacing, journald logs, and economy state byte-identical throughout (integrity ok, fk 0).
30 lines
1.1 KiB
Desktop File
30 lines
1.1 KiB
Desktop File
[Unit]
|
|
Description=OpenFUT: publish the FIFA17 container network namespace to /run/netns
|
|
Documentation=file:///home/alex/OpenFUT/scripts/systemd/openfut-netns-bind.sh
|
|
# PRODUCTION TEMPLATE — NOT INSTALLED. Staging runs in the host netns and needs
|
|
# none of this; only production enters the `openfut-fut-backend` container's
|
|
# namespace.
|
|
After=docker.service
|
|
Requires=docker.service
|
|
|
|
[Service]
|
|
Type=oneshot
|
|
RemainAfterExit=yes
|
|
|
|
# Re-resolves the container by NAME on every start, so no pid is ever hardcoded
|
|
# and a container restart is repaired by restarting this unit. It exits non-zero
|
|
# when the container is not running, which is what makes Core's Requires= on it
|
|
# a real admission gate rather than decoration.
|
|
ExecStart=/home/alex/OpenFUT/scripts/systemd/openfut-netns-bind.sh openfut-fut-backend openfut
|
|
|
|
# Deliberately NO ExecStop unmount. Tearing the bind mount down while Core and
|
|
# the host are still inside that namespace would strand them; the namespace is
|
|
# owned by the container's lifetime, not by this unit's.
|
|
|
|
StandardOutput=journal
|
|
StandardError=journal
|
|
SyslogIdentifier=openfut-netns
|
|
|
|
[Install]
|
|
WantedBy=multi-user.target
|