ops(systemd): follow the anchor container's netns across recreation
Recreating the Docker anchor left the supervised Core and host stranded in the dead namespace while systemd still reported them active — serving nobody, invisible to any monitoring that trusts unit state. Reproduced on staging: netns 4026539938 -> 4026540033, both pids unchanged in the old one, both units "active", traffic ConnectionResetError. There is no systemd-native edge signal to bind to. Containers do appear as units, but the scope name embeds the container ID (docker-<id>.scope), which changes on every recreate, so BindsTo= has no stable target; NetworkNamespacePath= resolves once at start; a .path unit on /run/netns would watch the file this tooling maintains. So: a level-triggered reconcile on a 10s timer, comparing the namespace the services are ACTUALLY in against the anchor's CURRENT one, acting only on a real difference. That cannot miss an event while the watcher restarts or dockerd is down, and needs no debounce — a burst of three recreations produced exactly one rebind. The trigger stays separable: a docker-events unit could invoke the same script. Anchor absent stops the dependants rather than falling back to host networking; docker unavailable logs once and retries on the next tick. Two defects found while testing and fixed here: - mount --bind STACKS when the old mount is busy, silently leaking nsfs entries; the bind helper now drains stale mounts in a loop. - reconcile must stop -> rebind -> start, not rebind -> restart: a running service holds the old namespace open and makes the umount fail busy. Staging also gained a faithful anchor container so the reproduction is structural rather than mocked. Economy state was byte-identical across every lifecycle test. Production units are templates only and remain uninstalled.
This commit is contained in:
@@ -49,7 +49,22 @@ if mountpoint -q "$TARGET" 2>/dev/null; then
|
||||
exit 0
|
||||
fi
|
||||
echo "openfut-netns-bind: $TARGET is STALE ($HAVE, want $WANT) — refreshing"
|
||||
umount "$TARGET" || true
|
||||
# DRAIN, do not just pop. `mount --bind` STACKS: binding over a busy mount
|
||||
# silently leaves the old one underneath, and staging grew two nsfs entries
|
||||
# on the first rebind before this loop existed. Left alone that is one
|
||||
# leaked mount per container recreation, and the buried namespaces are
|
||||
# exactly the dead ones we are trying to get rid of.
|
||||
#
|
||||
# A umount can legitimately fail while a service still holds the old
|
||||
# namespace open; the caller's job is to stop dependants FIRST. If it is
|
||||
# still busy we stack rather than fail — a current top-of-stack mount is
|
||||
# correct, just untidy — and say so.
|
||||
while mountpoint -q "$TARGET" 2>/dev/null; do
|
||||
umount "$TARGET" 2>/dev/null || {
|
||||
echo "openfut-netns-bind: $TARGET still busy; stacking a current mount over it" >&2
|
||||
break
|
||||
}
|
||||
done
|
||||
fi
|
||||
|
||||
[ -e "$TARGET" ] || touch "$TARGET"
|
||||
|
||||
Reference in New Issue
Block a user