fix(fifa17): complete kit stats, restore red squad tests, unrot prod gate

Four defects found by running the suites and the staging lifecycle end to end
after the kit milestone.

1. club-stats kits were half-implemented. The global `kits` counter was real
   but `kitsHome`/`kitsAway` and every per-team `kits` bucket stayed hardcoded
   0, so the same screen reported two owned kits and zero home/away kits.
   `kits` is a total with a family split, exactly like players/playersGold and
   staff/staffManager. The split key is `fcc_kitcards.assetid`: 14 is the home
   family and 15 the away family, verified across all 1482 rows of the kit
   table (assetid 14 covers exactly the 63xxxxx carddbids, 828 rows; assetid 15
   exactly the 64xxxxx ones, 654 rows; no exceptions either way).
   ClubStatInput now carries `asset_id`, and a kit buckets onto the team that
   wears it -- including a team the club owns no player from, the normal case
   for a kit won from a pack. The host reads both from the catalog through new
   NON-MINTING accessors: `resolve`/`resolve_kit` allocate a wire id, which a
   read-only stats query must never do as a side effect.

2. host_test.rs had 10 tests red since the squad-manager work (25f4ad1 /
   d37a9d5); 56bd9dd updated the squad_projection integration test and stopped
   there. `put_body` hardcoded the captured manager ref 100000427 into EVERY
   save, including tests with no manager fixture, so each one was refused with
   `unresolved_wire_ids` -- the tests were reporting a real invariant against a
   fixture that could not satisfy it. The manager is now an explicit
   `Option<i64>` per test, and FakeCore models Core's manager persistence
   instead of inheriting the "not implemented" default that 502'd every save.
   Added the coverage whose absence let this rot: a manager assignment
   round-trips as a Core owned id, a later save without one CLEARS it, and an
   unowned manager ref refuses the whole save with nothing committed.

3. `club_route_maps_query_and_shapes_core_items` pinned `offset`/`limit`
   forwarding to Core, which the kit commit deliberately replaced with
   host-side pagination. It only ever passed because FakeCore ignored the
   window -- against a real Core, `start=10` over a one-item club was always an
   empty page. Retargeted to the real contract (Core gets semantic filters and
   NO window) plus a new test that the window is applied locally after
   filtering, which the old fake made vacuous.

4. The staging lifecycle scripts identified production by hardcoded pids, so a
   correct teardown FATAL'd: production moved into containers and pids
   3631953/3374264 died with a container restart days ago. A pinned pid rots
   into the worst of both worlds -- a kill-refusal gate that no longer names
   any real production process, and a liveness gate that fails a healthy
   teardown. New shared `scripts/openfut_production.py` resolves production
   pids AND published ports from the container runtime at the moment they are
   needed, refuses to signal anything it cannot see, and proves production is
   the same processes serving the same ports before and after. Both lifecycle
   scripts use it, which also closed a real gap: port 8085 is published by
   openfut-fut-backend but was missing from the up script's forbidden list, so
   staging could have bound a production port.

Also fixes the economy differential, red because `complete_match` unlocks
achievements in the same transaction that pays the match reward -- a deliberate
Core feature the Python oracle has no counterpart for. `rust WIN +400` asserted
that progression did not exist; it now asserts the delta is the 400 match reward
plus exactly the achievements the match unlocked, read from Core's own report.
This commit is contained in:
funman300
2026-08-21 04:10:02 +00:00
parent db743ffd1f
commit 3442eac6f0
7 changed files with 568 additions and 107 deletions
+31 -46
View File
@@ -10,10 +10,12 @@ ones. So there is no pattern matching here at all:
* before any signal, /proc/<pid>/cmdline is read and MUST contain the staging
directory -- production's cmdline never can, because staging runs binaries
copied into that directory;
* the known production pids are refused explicitly, as a second gate;
* production's CURRENT pids, resolved from the container runtime, are refused
explicitly as a second gate;
* only the process GROUP the up script created (pgid == pid, via
start_new_session) is signalled, so a responder thread/child cannot be orphaned;
* afterwards every staging port is proven free and production is proven alive.
* afterwards every staging port is proven free, and production is proven to be
the same running containers serving the same ports as before.
python3 scripts/sold-staging-down.py
python3 scripts/sold-staging-down.py --purge # also delete the staging dir
@@ -29,20 +31,20 @@ import signal
import sys
import time
sys.path.insert(0, os.path.dirname(os.path.abspath(__file__)))
from openfut_production import ( # noqa: E402
PROD_PORTS,
ProductionError,
listening_ports,
production_state,
)
DEFAULT_STAGING_DIR = "/home/alex/openfut-sold-staging"
FORBIDDEN_PATHS = ("/home/alex/openfut-promotion/state",)
# Production processes that MUST be alive before and after this script runs. These
# two are the ones the batch contract names, and they live in the host pid view.
PROD_PIDS = {3631953: "prod utas-host", 3374264: "prod Core"}
# Reported but not gated: container pids change when the operator restarts the
# container, and a stale entry here would turn a successful teardown into a FATAL.
PROD_PIDS_INFO = {2090886: "prod blaze", 2091170: "prod python oracle",
2090888: "prod pow"}
PROD_PORTS = (8099, 8199, 18080, 8443, 42127, 42130, 42131, 4216, 8080, 8081, 8094)
class Fatal(RuntimeError):
class Fatal(ProductionError):
pass
@@ -87,37 +89,17 @@ def cmdline_of(pid: int) -> str:
return ""
def listening_ports() -> set[int]:
"""Ports in state LISTEN, from the kernel socket table. A trial bind() would
report EADDRINUSE for a stopped server's TIME_WAIT sockets and wrongly claim the
teardown failed."""
ports: set[int] = set()
for path in ("/proc/net/tcp", "/proc/net/tcp6"):
try:
with open(path) as fh:
next(fh, None) # header
for line in fh:
fields = line.split()
if len(fields) < 4 or fields[3] != "0A": # TCP_LISTEN
continue
ports.add(int(fields[1].rsplit(":", 1)[1], 16))
except OSError:
continue
return ports
def port_free(port: int) -> bool:
return port not in listening_ports()
def stop_one(rec: dict, staging_dir: str) -> str:
def stop_one(rec: dict, staging_dir: str, prod_pids: dict[int, str]) -> str:
"""Stop exactly one recorded process. Returns a human-readable outcome."""
name, pid = rec["name"], int(rec["pid"])
known_prod = {**PROD_PIDS, **PROD_PIDS_INFO}
if pid in known_prod:
if pid in prod_pids:
raise Fatal(
f"manifest entry {name} names PRODUCTION pid {pid} ({known_prod[pid]}). "
f"manifest entry {name} names PRODUCTION pid {pid} ({prod_pids[pid]}). "
"REFUSING to signal anything from this manifest."
)
if not pid_alive(pid):
@@ -192,8 +174,14 @@ def main() -> int:
)
step(f"variant : {manifest.get('variant')}")
# Resolved BEFORE anything is signalled: the refusal gate below is only
# meaningful if it knows production's pids as they are right now.
before = production_state()
for line in before.describe():
step(f"production : {line}")
for rec in manifest.get("processes", []):
ok(stop_one(rec, staging_dir))
ok(stop_one(rec, staging_dir, before.pids))
banner("PROVE STAGING IS GONE")
ports = manifest.get("ports", {})
@@ -214,16 +202,13 @@ def main() -> int:
ok("no recorded staging process is alive")
banner("PROVE PRODUCTION IS STILL UP")
dead = [f"{what} pid {pid}" for pid, what in PROD_PIDS.items()
if not pid_alive(pid)]
for pid, what in PROD_PIDS.items():
if pid_alive(pid):
ok(f"{what} pid {pid} alive")
if dead:
raise Fatal("production process(es) NOT alive: " + ", ".join(dead))
for pid, what in PROD_PIDS_INFO.items():
state = "alive" if pid_alive(pid) else "not found (informational only)"
step(f"{what} pid {pid} {state}")
after = production_state()
for line in after.describe():
ok(f"{line} alive")
after.assert_unchanged(before)
after.assert_serving()
ok(f"all {len(after.published)} published production ports still listening: "
+ ", ".join(str(port) for port in sorted(after.published)))
if args.purge:
shutil.rmtree(safe_path(staging_dir), ignore_errors=True)
@@ -243,7 +228,7 @@ def main() -> int:
print()
print(" then RELAUNCH the FIFA 17 client. See docs/SOLD_STAGING_RUNBOOK.md.")
return 0
except Fatal as exc:
except ProductionError as exc:
print(f"\nFATAL: {exc}", file=sys.stderr)
return 1