f55c401b6ceb09997d87da08993f4dc3b3e18f46
8 Commits
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
04c5043aba |
utas: My Squad filter corpus — every filter identified, root cause measured
Controlled retail capture, one criterion at a time, cleared between each. 47 transactions. Every filter the My Squad picker sends is now known from the wire rather than guessed. Route: GET /ut/game/fifa17/club -- the picker hits UTAS and reuses the general club-inventory route. level=any|gold quality lowercase, ALWAYS present rare=SP "Special" uppercase, OMITTED when off position=ST position uppercase, omitted when off nation=52 entity id numeric league=13 entity id numeric team=5 entity id numeric, NESTED under league sort=desc client constant; the UI has no sort control start=/count=11 pagination Two encoding families: short string enums, and numeric FIFA ids. The ids must never reach Core. Filters compose as plain ANDs in one query -- string and id filters alike -- so each maps independently. THE ROOT CAUSE IS SELF-AMPLIFYING. club_route honours type, team and league; it never reads start, count, level, sort or year. Because start is ignored, every page returns the same full set, so the client concludes the page was full and asks for the next one. One scroll produced 22 requests and 6.2 MB, stopping at start=200 only because the client gave up -- against a filtered set of 32 items that should have been three pages. That also explains why the bug reads as erratic rather than broken: league=13&position=ST returns every Premier League player instead of Premier League strikers. Plausible, wrongly sized, hard to notice. Measured filtered sets, from the real cluttered club -- these are the acceptance test for the fix: unfiltered 1962 league=13 350 league=13&team=5 32 Two client behaviours worth carrying forward: the picker fires a query per highlighted entry, not per selection (two requests for one club pick), and parameter ORDER is not stable, so parsing must be key-value. FIXTURE SIZE: bodies over 4 KB are truncated in the committed fixture, with body_full_len and body_full_sha256 retained, because the same 1.1 MB club response repeats ~25 times and its hash already proves identity. 13.3 MB -> 247 KB. The raw .ofcap keeps every byte, privately and gitignored. Truncation is recorded per transaction so a trimmed fixture is never mistaken for a whole response. Audited across all three identifier surfaces -- headers, JSON bodies, query strings -- before and after the size change: no leaks. 6/6 sanitiser mutations still killed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
0b66662525 |
utas: first real corpus, and two sanitiser gaps the audit caught
24 transactions across 11 connections from a retail session: login, hub, one pack open, two squad saves, a quick-sell, with before/after state manifests. Raw .ofcap stays gitignored at 0600; the sanitized corpus is committed as adapter fixtures. TWO GAPS FOUND BY AUDITING THE OUTPUT, NOT BY TRUSTING THE SANITISER. 1. `POST /ut/auth` carries `macAddress` and `deviceId`. Session tokens were being redacted correctly and these were not. A committed fixture is a published fixture. 2. Then, with those fixed, the audit fired AGAIN on the file about to be committed: `GET .../phishing/trusteddevice?deviceId=...` puts the id in the QUERY STRING. Three input surfaces carry identifiers -- headers, JSON bodies, and query strings -- and the sanitiser knew about two. Both fixed in the tool rather than by editing the file, with a regression test and a mutation for the query path. AND A THIRD ARTEFACT MIX-UP, in the mutation harness itself. It reported the query-redaction mutation as SURVIVED while a hand-run of the same mutation killed it. Cause: the harness pointed at a stale scratchpad copy of the test that pre-dated the query assertion, so it was faithfully testing the mutated tool against a test that could not detect the mutation. That is the same class as the build guard checking the wrong binary and cargo reusing a binary compiled from mutated source -- the third instance today of measuring the wrong artifact. The harness now resolves ROOT from its own location and runs the COMMITTED test; the stale copy is deleted. Harness committed as scripts/mutate-utas-observe.py so this is repeatable rather than a thing that happened once in a scratch directory. 6/6 killed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
cdea85e214 |
utas: standalone recording proxy that tees rather than rebuilds
UTAS needs a real request/response corpus before any Rust is written: it is where protocol shape and FUT state start being coupled, so guessing is worse here than it was for Blaze. The oracle truncates logged bodies at ~200 chars, and raising that cap would mean editing the behavioural specification to make it easier to copy -- backwards. A proxy gets the same evidence and leaves the oracle untouched. THE DESIGN RULE: TEE, DO NOT REBUILD. UTAS is plaintext HTTP/1.1 on ThreadingHTTPServer, so keep-alive, pipelining and chunked transfer are all live. A proxy that parses a request and re-emits it can corrupt the traffic it exists to observe -- and that corruption would present as a UTAS bug, pointing the investigation in exactly the wrong direction. So bytes are copied verbatim in both directions and a second copy goes to disk; transactions are reconstructed later, offline, from that copy. A parser bug therefore spoils the record and never the session. Standalone, NOT in the container, so the same tool can later sit in front of a Rust UTAS host and replay an identical captured request against both. Two layers, as with the Blaze captures: raw/*.ofcap is exact bytes at mode 0600 and gitignored; sanitized/transactions.jsonl is the committed artefact. Bodies are preserved EXACTLY and sanitised second -- only known-secret headers and JSON keys are replaced, structure is never reshaped, and every redaction is recorded in the transaction so a reader knows what was touched. Captured per transaction: connection id, sequence, relative and wall time, elapsed ms, method, path, query, HTTP version, headers IN RECEIVED ORDER as pairs (a dict would drop duplicates and ordering), raw body and length for both directions, status, and observed keep-alive. Verified as two independent properties, because they fail differently: transparency (bytes through the proxy identical to bytes direct, Date masked, with the mask asserted to have fired) and fidelity (parsed transactions match what was sent, including a dechunked response and a 300-byte POST body). 5/5 mutations killed, including "record but do not forward", "drop the last byte of every chunk" and "stop redacting". scripts/test-utas-observe.py is committed alongside it: a capture tool nobody can re-verify is not evidence. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> |
||
|
|
8f3b659c33 |
lifecycle: one host-lifecycle helper; roster.sh; ban pkill -f
redirector.sh and the coming roster.sh needed the same five rules, each
of which cost something to learn:
* resolve /proc/PID/exe; never match a command line. `pkill -f` /
`pgrep -f` match any shell whose ARGUMENTS mention the name, including
the shell running the command. That has killed this session's own
shell twice, and is now banned in migration tooling -- the helper
contains no `-f` matching and the header says why.
* `readlink`, not `readlink -f`. After a rebuild the link reads
"<path> (deleted)" and -f resolves it to nothing, so the orphan check
goes blind to exactly the long-lived processes it exists to find. Two
orphans hid there, one serving the wrong certificate.
* stop PROVES the process is gone and the port free.
* an ambiguous binary is an error for start/verify but NOT for
stop/status: rollback must never be blocked by a question about the
build tree.
* verify the RUNNING process's commit, not the artifact on disk, which
a rebuild can silently advance past.
Copying those into a second script would have been the same mistake as
copying the TLS setup. Instead scripts/host-lifecycle.sh owns them and a
service supplies four facts: name, crate, executable, port variable.
redirector.sh goes from 178 lines to 26 and roster.sh is 24, with no
behaviour change -- the refactored redirector.sh still sees the live
armed process (pid 830736, port 42227) and still refuses correctly
because HEAD has moved past it.
Paths are unchanged (rundir, pidfile, portfile, commit stamp, log), so
the currently running redirector stays manageable across this refactor.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
|
||
|
|
fc411bb6f1 |
scripts: require one certificate across the whole FIFA-facing TLS stack
Two live gates were lost to a second variable I had been asked to eliminate. The Rust redirector was pointed at the repo's fifa17-recon/tools/redir_cert.pem (fingerprint F9:16:1A...), while the running container serves a different cert baked into its image (E7:F9:46...) which the Python redirector, roster and the rest of the stack all share. So the A/B compared TLS implementation AND certificate identity at once. FIFA 17's ProtoSSL caches the server certificate for a backend. The redirector is the first TLS connection of a session, so its cert becomes the one the client expects; the next service presenting a different cert fails its handshake. That is why the redirect itself always succeeded and the failure surfaced later, on the roster fetch -- "An error occurred downloading the FUT Squad Update". It stayed invisible because Python's socketserver swallows it: a handshake failure at accept() raises ssl.SSLError, which subclasses OSError and is discarded by _handle_request_noblock. No request log, no stderr. Every server looked healthy while the client could not talk to any of them. Confirmed on the wire: tls-observe in front of the roster server captured four ClientHellos from the client, correct SNI and the same 8 static-RSA suites it offers the redirector, none of which produced a request. The check is mutation-tested against the real bug: with a redirector started on the stale repo cert it exits 1 and names the mismatch. |
||
|
|
5bc39e902d |
tooling: observe client connection ATTEMPTS; make the build guard reject bad args
openfut-observe.sh answers the one question no server log can: when a gate fails and a service logged nothing, did the client try and fail, or never try? Both look like silence. Two redirector gates were lost to that ambiguity -- "roster server logged nothing" was equally consistent with a broken roster service, a wrong roster URL, and a client that never asked. Built on iptables packet counters because this box has no tcpdump, no conntrack, and no readable kernel log. That last one is verified rather than assumed: an initial LOG-based version installed correctly and its rules matched (counters proved it), but the output went nowhere -- journalctl -k has no entries and dmesg is empty. Counters are also lower volume and record only SYNs, so no payload can be captured even in principle. Validated against the live client, not a loopback stand-in: an initial self-test using this host's own address counted almost nothing, because locally-generated packets never traverse PREROUTING. Against the real remote client it counts 8081 at ~4/min, matching the roster server's own log. Known gap, recorded rather than hidden: the catch-all TOTAL runs well above the sum of the named ports, so the client makes steady background attempts to ports not tracked here. It is present during a working session, so it is not the failure signature, and it is not chased further here. verify-build-identity.sh now rejects an argument that is not a commit hash. Passing the binary path instead of its stamp previously produced a plausible "REFUSING: binary was built from ./target/release/... but HEAD is <sha>", which reads as a real stale-build finding rather than a caller mistake -- and a safeguard that cries wolf is one people learn to route around. Usage error is now exit 2, distinct from a genuine stale build (1) and success (0). |
||
|
|
c03702707b |
redirector: commit stamp + shared build-identity verifier that REFUSES
The binary records only the commit it was built from -- no dirty-tree flag. Cargo will not re-run a build script because another crate's source changed, so a compiled-in 'clean' claim can be stale and is not a safeguard; that was verified on the Blaze host. scripts/verify-build-identity.sh establishes both facts at LAUNCH, where they cannot go stale: the stamped commit equals HEAD, and the migration crates are clean. It REFUSES rather than warns, because for a migration gate a warning on stderr is something to scroll past. --identity prints the stamp without valid configuration. The launcher must be able to establish which commit a binary came from BEFORE deciding whether to run it; requiring a correct environment first would invert the check. redirector.sh mirrors sidecar.sh: refuses to start with an orphan present or the port busy, matches the resolved executable rather than the command line (pgrep -f matches any shell mentioning the name), and stop PROVES the process is gone and the port free. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> |
||
|
|
8f5f54833f |
ci: tripwire against lab addresses creeping back into tracked source
Cheap insurance, explicitly not the real check -- the semantic tests in
deployment_config.rs are what prove propagation, using two TEST-NET addresses
and bind != advertise. This grep only stops the lab subnet reappearing months
from now when the reasoning has been forgotten.
Deployment config legitimately contains real addresses and lives in gitignored
files, so it is never scanned. The frozen baseline doc is allowlisted BY PATH:
it records what a past deployment actually was, and rewriting it would falsify
the record.
Also swapped the lab IP for a TEST-NET placeholder in the usage examples and
error messages of compose/entrypoint/client_arm. Those were already correct
architecture -- every one requires the address via ${VAR:?} -- but using the
real lab IP as the example is the same 'happens to match our lab' smell, and
placeholders keep the tripwire allowlist near-empty.
Mutation-tested: adding a lab address to a source file makes it exit 1.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|