# OpenFUT systemd units Supervision for OpenFUT Core and the FIFA17 UTAS host. Replaces the previous `setsid nohup … nsenter …` launch, which had no restart policy, no boot persistence and no supervisor-visible logs. | file | scope | installed? | |---|---|---| | `openfut-staging-core.service` | staging Core, port 18081 | **yes** — proving ground | | `openfut-staging-host.service` | staging host, port 8299 | **yes** | | `openfut-netns.service` | publishes the container netns to `/run/netns/openfut` | template only | | `openfut-core.service` | production Core, port 18080 | template only | | `openfut-host.service` | production host, port 8099 | template only | | `openfut-netns-bind.sh` | resolves the container netns by NAME, idempotently | helper | | `openfut-wait-tcp.sh` | bounded readiness gate | helper | Production templates are **not installed**. Deploy only via the plan in `OpenFUT-Vault/06 Operations/OpenFUT Service Supervision (staging-proven).md`. ## The three decisions worth knowing **1. `Wants=`, not `Requires=`, from host → Core.** Measured on staging: `Requires` propagates a Core stop into a host stop, and a later Core start does *not* bring the host back — a routine Core restart would leave the client with no server. With `Wants`, the host survives a Core outage, answers `503 core_unavailable` (never a Python fallback), and resumes the moment Core returns, with no supervisor intervention. **2. Readiness is an `ExecStartPre` TCP gate, not ordering.** `After=`/`Wants=` order units; `Type=exec` only proves the binary exec'd. Neither means Core can serve. Core binds its listener *after* opening the DB, running migrations and loading the content pack, so "port open" is a genuine readiness signal. The gate is bounded and *fails* rather than blocking: a host that waits forever looks healthy to the supervisor while serving nobody. **3. The netns is resolved by container NAME at every start.** Production must run inside `openfut-fut-backend`'s network namespace. The container is `restart=unless-stopped`, and its netns inode *changes* on restart — observed `net:[4026539938] → net:[4026540033]`. A hardcoded pid is therefore wrong by construction, and any process left in the old namespace keeps running with no interfaces, silently serving nobody. `openfut-netns-bind.sh` re-resolves by name and refreshes a stale bind mount; `NetworkNamespacePath=` then enters it declaratively. > **Still open:** a container restart strands *already-running* services in the > dead namespace. Re-running `openfut-netns.service` plus restarting Core/host > repairs it, but nothing triggers that automatically yet. See the promotion > plan's "Residual gap". ## Operating ```bash systemctl status openfut-staging-core openfut-staging-host journalctl -u openfut-staging-core -f sudo systemctl restart openfut-staging-core # host survives and recovers sudo systemctl stop openfut-staging-host openfut-staging-core ``` Config lives in `EnvironmentFile`s (`…/systemd/core.env`, `host.env`), generated from the live process environment so supervision changed *how* the processes start and nothing about *what* they do. Binaries are immutable copies, so a later `cargo build` cannot change what is running. ## Interaction with the staging lifecycle script `scripts/sold-staging-up.py` still starts its own unsupervised processes and refuses to run while a staging stack is up. Stop the units first: ```bash sudo systemctl stop openfut-staging-host openfut-staging-core python3 scripts/sold-staging-up.py --club real --variant highest … ``` Reconciling the two (having the script drive the units) is deliberately out of scope for the supervision milestone.