Files
OpenFUT/fifa17-recon/tools/login_dump/QOS_CONNECT_PLAN.md
T
funman300 edab23f04a fifa17-recon: package the working offline FUT backend
Emulates FIFA 17's full online + Ultimate Team stack against an offline,
clean-room backend (no EA servers). Proven end-to-end 2026-08-01:
Origin login -> Blaze login -> device-trust -> the FUT hub.

Package:
- tools/openfut-fut.sh   one-command orchestrator (start/stop/status/restart)
- tools/root_arm.sh      idempotent host arm (sysctls, DNAT, /etc/hosts easw)
- tools/{lsx_responder_v2,blaze_responder_v3b,roster_server,utas_server,autopatch}.py
  the 5 servers (Origin LSX :4216, Blaze :42127/42130/42131, roster :8081,
  FUT/UTAS :8099) + heat2.py (Fire2/Heat2 TDF codec)
- FUT-RUNBOOK.md         runbook + gate-ladder troubleshooting
- docs/, tools/login_dump/*.md   the reverse-engineering write-ups

All findings are clean-room, from binaries we own; nothing from any leak.
The wire protocol maps 1:1 to FIFA 23.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PN5bmpDVQR1aXgefyWAt7o
2026-08-01 09:12:17 -07:00

29 KiB
Raw Blame History

QOS_CONNECT_PLAN — "Unable to connect to the EA servers"

Date: 2026-07-31 · Target: FIFA17.exe (decrypted, VMA==runtime VA), live pid 50676 Sources: decrypted image + /proc/50676/mem + /tmp/blaze_responder.log + /tmp/lsx.log + blaze_responder_v3b.py. Clean-room: no leaked EA source was consulted.


0. Verdict up front

The QoS hypothesis in BRIEF5 is WRONG. The empty LTPS and the unserved 127.0.0.1:17502 are both harmless and were never even reached. The bug is a silent type error in our CONF map: three of the values we send are TimeValue-typed and FIFA's duration parser rejects bare integers, so pingPeriod, defaultRequestTimeout and connIdleTimeout all land as 0. The one-line class of fix is "30000000""30s".

BRIEF5's premise is also factually wrong and should be retired: session 1 did not stay up for ~2.5 minutes of pings. It died at 16:38:07, one second after the login burst, and the whole log contains exactly one Util::ping ever. The "2.5 minutes" is dead air between two failed connections.

Three of the four independent reverser passes converged on this cause. Everything load-bearing below was re-verified from scratch for this synthesis (live reads + raw bytes re-disassembled), and two new pieces of proof were added that no individual report had — see §1.5 and §1.6.


1. THE CAUSE — end to end

1.1 What we actually send

blaze_responder_v3b.py:400-406, inside PreAuthResponse.CONF (a map<string,string>):

("autoReconnectEnabled",  "1"),          # OK  — read via atoi
("connIdleTimeout",       "90000000"),   # BROKEN — TimeValue
("defaultRequestTimeout", "30000000"),   # BROKEN — TimeValue
("maxReconnectAttempts",  "5"),          # OK  — read via atoi
("pingPeriod",            "20000000"),   # BROKEN — TimeValue

We wrote microseconds as bare integers. FIFA reads these five keys through two different vtable getters, and only one of them accepts a bare integer.

1.2 The two getters

ConnectionManager vtable (live vptr 0x1438a0850):

Slot Function Semantics
vt+0x48 0x146e1bb80-adjacent raw getConfigString
vt+0x50 0x146e1bb80 uint32 — string → atoi (import 0x148e21ad8). Works.
vt+0x58 0x146e1bda0 TimeValue — string → duration parser 0x1479b2d50. Broken for us.

0x146e1bda0 re-disassembled live (raw bytes read out of pid 50676 for this document):

146e1bda0: 53                    push rbx
146e1bda1: 4883ec20              sub  rsp,0x20
146e1bda5: 488b01                mov  rax,[rcx]
146e1bda8: 4c89c3                mov  rbx,r8
146e1bdab: 4c8d442430            lea  r8,[rsp+0x30]
146e1bdb0: ff5048                call [rax+0x48]        ; getConfigString
146e1bdb3: 84c0                  test al,al
146e1bdb5: 7506                  jne  0x146e1bdbd
146e1bdb7: 4883c420 5b c3        add rsp,0x20; pop rbx; ret     ; key absent -> false
146e1bdbd: 488b542430            mov  rdx,[rsp+0x30]
146e1bdc2: 4889d9                mov  rcx,rbx
146e1bdc5: e8866fb900            call 0x1479b2d50       ; TimeValue::parse   <-- return value
146e1bdca: b001                  mov  al,0x1            ; <-- CLOBBERED, UNCONDITIONAL "success"
146e1bdcc: 4883c420 5b c3        add rsp,0x20; pop rbx; ret

Line 146e1bdca: b0 01 is the whole problem: the parser's boolean result is thrown away. The caller is told the value was applied. It was not.

1.3 The parser rejects us

0x1479b2d50 accumulates digits, then dispatches on the byte that terminates the run. Raw bytes at 0x1479b2db4, read live:

1479b2db4: 80f964 7435    cmp cl,'d' ; je
1479b2db9: 80f968 742b    cmp cl,'h' ; je
1479b2dbe: 80f96d 7414    cmp cl,'m' ; je      (+ 's' lookahead -> ms)
1479b2dc3: 80f973 740a    cmp cl,'s' ; je
1479b2dc8: 80f979 7535    cmp cl,'y' ; jne 0x1479b2e02
...
1479b2e02: 30c0           xor al,al            ; return false
1479b2e04: eb50           jmp 0x1479b2e56      ; SKIPS the store at 0x1479b2e46

Accepted grammar: optional -, then :-separated <digits><unit> with unit ∈ {y, d, h, m, ms, s}. A NUL terminator with no unit — i.e. exactly "20000000" — falls straight to 0x1479b2e02, returns false, and never executes mov [r14],rax. The caller pre-zeroes each out-slot (0x146e1d0af, 0x146e1d10c, 0x146e1d143, all mov QWORD PTR [rbp+N],rdi with rdi=0), so the value the SDK ends up applying is 0.

1.4 Live proof that all three landed as zero

Read out of pid 50676 for this document. ConnectionManager = 0x43c47ff0 (vptr 0x1438a0850 ✓, CM+0x1288 = "Blaze 15.1.1.3.0 (OpenFUT)\n" — our own SVER, so this is the object our reply configured):

Field Offset Written at Live value What we intended
pingPeriod CM+0xd1c 0x146e1d0c9 / 0x146e1d0f5 15000 20000
defaultRequestTimeout CM+0x278 0x146e1d12c 0 30000
connIdleTimeout CM+0xd28 0x146e1d164 0 90000
maxReconnectAttempts CM+0xc5c 0x146e1d1b6 5 ✓ 5
autoReconnectEnabled CM+0x11b7 0x146e1d188 1 ✓ 1

The split is exactly along the getter boundary: every atoi key took, every TimeValue key was dropped. That is not a coincidence, it is the signature.

The 15000 is not a "key missing" default — it is the clamp. Bytes at 0x146e1d0a5, live:

146e1d0a5: 488d1584 39a8fc   lea rdx,[rip-0x357867c]  ; -> 0x1438a0a30 = "pingPeriod"
146e1d0ac: 4889d9            mov rcx,rbx
146e1d0af: 48897d07          mov [rbp+7],rdi          ; pre-zero out-slot
146e1d0b3: ff5058            call [rax+0x58]          ; getConfigTimeValue  -> always true
146e1d0b6: 48be cff753e3a59bc420  movabs rsi,0x20c49ba5e353f7cf   ; /1000 magic
146e1d0c0: 84c0              test al,al
146e1d0c2: 750d              jne  0x146e1d0d1
146e1d0c4: b8983a0000        mov  eax,0x3a98          ; 15000 = key-absent default
146e1d0c9: 89831c0d0000      mov  [rbx+0xd1c],eax

al is always 1, so we take the jne — the parsed value (0) is divided by 1000 → 0, then the cmovl at 0x146e1d0f2 clamps anything below 1000 up to 0x3a98. Reaching 15000 through the clamp is positive proof the parse produced 0. The lea also confirms the key name byte-for-byte: 0x1438a0a30 is exactly the address our own source comment on line 406 cites.

1.5 NEW PROOF — session 2's response callback provably never ran

onPreAuthResponse (0x146e1cf10) has no branch on its success path that can skip 0x146e1d1cb: call 0x146e1e460 (sendPing). So "did a ping happen" is a perfect oracle for "did the callback run". The log shows exactly one ping ever (RX #3984, 16:38:06, conn 43427) and none in session 2.

onPingResponse (0x146e1d290) stamps CM+0x11a0 = response STIM and CM+0x11a4 = local tick. Live values, decoded for this document:

  • CM+0x11a0 = 17855410862026-07-31 16:38:06, and our responder builds STIM as int(time.time()) (blaze_responder_v3b.py:564) — this is literally the timestamp we generated and sent in session 1.
  • CM+0x11a4 = 14437471 ms. This is a system-monotonic tick, not process-relative. Against /proc/uptime = 16638.9 s at wallclock 17:14:48, it back-solves to 16:38:06.

Two independent clocks, both landing on 16:38:06, not 16:40:34. The preAuth response callback last completed in session 1 and did not run at all in session 2 — even though our reply was byte-identical (cmp -l rx_3983… rx_4009… differs in one byte, offset 13 = msgNum) and was fully received (the raw 772-byte payload is still sitting in the recv buffer at 0x43c4e2xx-0x43c4e373).

The teardown therefore precedes the dispatch of our reply. Nothing about the reply's content is being rejected. That kills the "QoS field validation" theory outright — the QoS fields are parsed inside the callback that never ran.

1.6 NEW PROOF — why session 1 survived preAuth and session 2 did not

This is the asymmetry BRIEF5 asked about, and it falls straight out of §1.4.

ConnectionManager::ctor at 0x146e187ea:

146e187ea: 488b5608     mov rdx,[rsi+0x8]      ; BlazeHub
146e187ee: 8b8244050000 mov eax,[rdx+0x544]
146e187f4: 89867802 0000 mov [rsi+0x278],eax   ; defaultRequestTimeout = hub+0x544

Live hub = 0x43c47330, hub+0x544 = 10000. So a freshly-constructed ConnectionManager has a healthy 10000 ms request timeout — and CM+0x278 is only ever overwritten inside the preAuth response callback (0x146e1d12c).

CM+0x278 while its own preAuth RPC was outstanding Outcome
conn A (43427, 16:38:06) 10000 ms (ctor default, not yet clobbered) preAuth callback ran → ping sent → login burst
conn C (35037, 16:40:34) 0 (clobbered by conn A's callback, persists for the CM's lifetime) preAuth job expires before the reply is dispatched → teardown

Our own preAuth reply poisons the timeout that the next preAuth needs. Session 1 worked because it ran before our config was applied. That is why the identical bytes produce opposite outcomes, and why no amount of varying the QOSS content would ever have helped.

The same zero explains the rest of the log: with connIdleTimeout = 0, every connection dies at its first idle moment — conn A right after the ping reply, conn B (43131) right after TX #4008.0, the last frame of the login burst, conn C right after our preAuth reply. Conn B's 24 RPCs all succeeded because localhost replies land inside the same SDK tick they were issued on, so the expiry sweep never saw them outstanding.

Conn B also explains itself: onSocketConnected (0x146e1cb60) gates preAuth on a flag —

146e1cbed: 4038bb3c0b0000  cmp BYTE PTR [rbx+0xb3c],dil    ; dil = 0 (errorCode)
146e1cbf4: 7507            jne 0x146e1cbfd
146e1cbf6: e8f2140000      call 0x146e1e0f0                ; sendPreAuthRequest
146e1cbfb: eb0c            jmp 0x146e1cc09
146e1cbfd: c68380110000 01 mov BYTE PTR [rbx+0x1180],1
146e1cc04: e854180000      call 0x146e1e460                ; sendPing only

Conn B skipped preAuth (its first frame is fetchClientConfig msgNum=2) → CM+0xb3c was non-zero → conn B was an SDK auto-reconnect (autoReconnectEnabled=1, maxReconnectAttempts=5, both live ✓). CM+0xb3c reads 0 now, so the flag is cleared on full teardown — consistent with conn C re-sending preAuth. And msgNum is one monotonic counter 0…26 across all three sockets, with live CM+0xc50 = 26 ✓ — there is only one ConnectionManager, so "session 2 is a different kind of connection" is false.

1.7 The QoS path is exonerated, with disasm

QosManager::startLatencyProbes 0x146e1dae0, re-disassembled live:

146e1dae0: 53                push rbx
146e1dae1: 4883ec20          sub  rsp,0x20
146e1dae5: 488b8120020000    mov  rax,[rcx+0x220]      ; LTPS end
146e1daec: 4889cb            mov  rbx,rcx
146e1daef: 48398118020000    cmp  [rcx+0x218],rax      ; LTPS begin
146e1daf6: 7508              jne  0x146e1db00
146e1daf8: 30c0              xor  al,al                ; EMPTY -> false, benign, no side effects
146e1dafa: 4883c420 5b c3    add rsp,0x20; pop rbx; ret

Its caller QosManager::initialize 0x146e1c3f0 has a dedicated no-probes-started branch:

146e1c5aa: e831150000        call 0x146e1dae0
146e1c5af: 84c0              test al,al
146e1c5b1: 7508              jne  0x146e1c5bb
146e1c5b3: 4889f9            mov  rcx,rdi
146e1c5b6: e875d5ffff        call 0x146e19b30          ; normal QoS-completion fn
146e1c5bb: c6471001          mov  BYTE PTR [rdi+0x10],1 ; initialized = true

Live QosManager = CM+0x1db0 = 0x43c49da0 (vptr 0x1438a07c8 ✓) proves that exact path ran to completion:

Field Live Meaning
Q+0x10 1 initialize() reached its final store — completed
Q+0x20 0 both QoS tests off — our enableQos*Test=false parsed fine
Q+0x218 / Q+0x220 0 / 0 LTPS empty → probe loop 0x146e1c4bb..0x146e1c4f0 never ran
Q+0x18 0 no DirtySDK QosApi object exists
Q+0x1c8 / Q+0x1f8 / Q+0x250 10 / 17502 / 5000000 our LNP / BWPS.PSP / TIME arrived intact
hub+0x53f 1 QoS is enabled; the early-out at 0x146e1c44b was not taken
CM+0x1f60 1 SDK's own "QoS ready" flag is set

So QoS ran, with QoS enabled, with an empty ping-site list, and reported ready.

Nothing ever contacted 17502. Verified for this document: no 445e in /proc/net/{tcp,udp,tcp6,udp6}, and ss -lntup shows our responder listening only on 42127 / 42130 / 42131. The BWPS port is merely the QosApi create argument (0x146e1975a: movzx r8d,WORD PTR [rbx+0x1f8]), and the real probe transport is HTTPS (https://%s:%u/qos/qos @ 0x143aa75c0) driven only by LTPS entries, of which there are none.

PinQosError_ReferenceEvent @0x1439de5c8 has zero code xrefs (it sits in a telemetry name table next to CM_WonTournament, PS3TOSAccepted, …) and POW:sConnectionManager @0x143995b88 is an allocation tag in the powdll loader string block. Both anchors in BRIEF5 are red herrings.

1.8 The on-screen message corroborates all of the above

The popup text is loc key OSDK_LOST_CON_TO_EA @0x14395c6c8 (resolved English on the heap at 0x7b438e0). Selection site 0x1471aa220, bytes re-read live at 0x1471aa286:

1471aa286: 83ba0c02000001  cmp   DWORD PTR [rdx+0x20c],1
1471aa28d: 488d0534247bfc  lea   rax,[rip-0x384dbcc]   ; 0x14395c6c8 OSDK_LOST_CON_TO_EA
1471aa294: 4c8d1d157d7afc  lea   r11,[rip-0x38582eb]   ; 0x143951fb0 OSDK_A_R30
1471aa29b: 4c0f44d8        cmove r11,rax

Caller 0x1471a1dda passes rdx = cnncMgr+0x4c4, so the tested field is cnncMgr+0x6d0 — the "an EA request is currently in flight" busy flag (set to 1 immediately before each async Blaze op at 0x1471b50bc, …5115, …5202, …5282, …5322, …5cb1, 0x1471b6578; cleared on completion at 0x14717d7bb, 0x147190d06, 0x1471b4fa4, 0x1471b501e, 0x1471b54a8, 0x1471b5c37, 0x1471b65c6, 0x1471d5486, 0x1471d6919).

So the message means, literally: "the socket died while an RPC was outstanding." Not "QoS validation failed", and not "login timed out" — the 30 s login timeout path (LoginStateVersionCheck::Tick 0x1471b4ee0, lea eax,[rdx+0x7530]) resolves to table index 0 and yields OSDK_A_R30, a different on-screen string. The alternative state-machine selector also did not fire: the live LoginStateLogin object (0x43d189d8, vtable 0x14395c180) still has +0x7c errCode = 0 and +0x80 = 0x143951fb0 (the ctor default), i.e. no error was ever latched.

The popup is therefore a pure transport-disconnect notice, and it is exactly what a request that expires under a 0 ms timeout would produce.

1.9 Causal chain, one paragraph

We send defaultRequestTimeout="30000000" in PreAuthResponse.CONF. FIFA reads it with the TimeValue getter 0x146e1bda0, whose parser 0x1479b2d50 requires a unit suffix and returns false on a bare integer; the getter discards that false (146e1bdca: b0 01) and the SDK applies the pre-zeroed out-slot, i.e. 0, to CM+0x278 — overwriting the ctor's healthy 10000 ms — plus CM+0xd28 = 0 (idle timeout) and CM+0xd1c clamped to 15000. This happens inside session 1's preAuth response callback, so session 1's own preAuth was still protected by 10000 ms and completed normally. From that instant on, every connection on this ConnectionManager dies at its first idle moment, and any RPC that is not answered within the issuing tick is expired immediately. When the user goes online at 16:40:34, conn C issues Util::preAuth under a 0 ms timeout; the job is expired before our reply is dispatched, the callback never runs (proved by the 16:38:06 ping tick), the SDK tears the socket down with cnncMgr+0x6d0 == 1, and 0x1471aa220 selects OSDK_LOST_CON_TO_EA: "Unable to connect to the EA servers at this time."


2. THE FIX — ranked and minimal

FIX 1 — REQUIRED, proven, 3 lines, no new listener or thread

blaze_responder_v3b.py:401-406. Change bare microsecond integers to unit-suffixed duration strings:

("connIdleTimeout",       "90s"),   # was "90000000" -> parsed to 0
("defaultRequestTimeout", "30s"),   # was "30000000" -> parsed to 0
("pingPeriod",            "20s"),   # was "20000000" -> clamped to 15000

"90000ms" / "30000ms" / "20000ms" are equally valid. Leave maxReconnectAttempts and autoReconnectEnabled as plain integers — they go through the atoi getter and already work.

Grammar accepted by 0x1479b2d50: optional leading -, then :-separated <digits><unit> with unit ∈ {y, d, h, m, ms, s}; combined as ((((y*365+d)*24+h)*60+m)*60+s)*1000+ms, then ×1000 → µs. Corroborating precedent from the shipped image: the default for QosConfigInfo.timeout is the literal string "5s" @ ~0x14388cc90.

While editing, add a comment recording the getter split so this class of bug cannot recur:

# CONF value TYPES matter.  Keys read via ConnMgr vt+0x50 (0x146e1bb80) go through
# atoi -> plain integers are fine.  Keys read via vt+0x58 (0x146e1bda0) go through
# TimeValue::parse (0x1479b2d50), which REQUIRES a unit suffix (y/d/h/m/ms/s) and
# silently yields 0 for a bare integer -- 0x146e1bdca discards the parser's bool.

Non-interactive verification (no relaunch needed to check the parse, only a reconnect): read the live ConnectionManager and assert the three fields flipped.

PID=$(pgrep -x FIFA17.exe)
python3 - <<'EOF'
import struct
PID=<pid>
f=open("/proc/%d/mem"%PID,"rb")
def u32(va): f.seek(va); return struct.unpack("<I",f.read(4))[0]
CM=0x43c47ff0                      # re-locate: scan rw memory for qword == 0x1438a0850
print("pingPeriod  CM+0xd1c =", u32(CM+0xd1c), "(want 20000, currently 15000)")
print("defReqTO    CM+0x278 =", u32(CM+0x278), "(want 30000, currently 0)")
print("connIdleTO  CM+0xd28 =", u32(CM+0xd28), "(want 90000, currently 0)")
EOF

Note CM=0x43c47ff0 is this process instance's address. After a relaunch, re-locate it by scanning rw regions for the 8-byte vptr 0x1438a0850 (exactly one hit), or for the string Blaze 15.1.1.3.0 (OpenFUT) and subtracting 0x1288.

Observable success signal (the one that matters): on the go-online connection, after our Util::preAuth reply the client emits Util::ping (component 0x0009, cmd 0x0002) instead of closed, and proceeds to Authentication::login (0x0001/0x000a) on that same socket. A secondary signal: the boot connection stops dying ~1 s after the login burst and starts emitting a ping every ~20 s.

FIX 2 — cheap, low risk, not the cause: close the CIDS gap

blaze_responder_v3b.py:521-525. We advertise CIDS = (1, 4, 5, 7, 9, 15, 25, 28, 30722), but the session-1 log shows the client talking to three components we never listed:

Component Seen in log Frame
0x000a (10) RX #3999 cmd:0x0005, RX #4008 cmd:0x0002 UserSessions-adjacent
0x000b (11) RX #4001 cmd:0x0a28
0x08c9 (2249) RX #4005 cmd:0x0001

Add 10, 11, 0x08c9 to COMPONENT_IDS. Not the cause (the client sent to them regardless and we replied), but it is free and it removes a known divergence from a real server. Do this after FIX 1 lands and only if FIX 1 alone does not clear the gate, so the two changes stay separable.

FIX 3 — DO NOT DO YET. Populating LTPS (and serving :17502)

Explicitly deferred, and it is the one item that would need a new listener and thread.

Reasons to leave qos_config() (blaze_responder_v3b.py:527-538) exactly as it is:

  1. The empty-LTPS path is a first-class, benign branch (0x146e1daf8: xor al,al) and live memory proves it completed with QoS enabled (Q+0x10=1, hub+0x53f=1, CM+0x1f60=1). It blocks nothing.
  2. The QoS manager is initialized at 0x146e1d214, after the sendPing at 0x146e1d1cb, inside a callback that session 2 provably never entered (§1.5). It is downstream of the failure.
  3. Populating LTPS creates work: the per-site loop 0x146e1c4bb..0x146e1c4f0 would then run and fire real HTTPS GET https://127.0.0.1:17502/qos/qos requests. Serving those needs a new TLS listener + thread in the responder, with a cert the client accepts. We would be trading a non-problem for a real one.
  4. The Heat2 framing for map<string, STRUCT> is flagged UNVERIFIED in heat2.py:204 (and the STRUCT value path in _enc_value likewise). A malformed TDF here would fail the client's decode of the whole PreAuthResponse — a strictly worse failure than the current one.

Only revisit if, after FIX 1, the game reaches login but then stalls on a ping-site-dependent feature. In that case the anchors are already mapped: QosConfigInfo copied to QosManager+0x1b8; +0x1c8=LNP, +0x1d0=BWPS (QosPingSiteInfo vtable 0x143889ba0), +0x200=LTPS map (vtable 0x143889c48, begin +0x218 / end +0x220, element stride 0x20, int32 latency at element+0x18 — stride and latency offset independently confirmed by the OSDK helper 0x1471c0c90, which loops add rax,0x20 writing mov DWORD PTR [rax+0x18],0x3e8). The minimal payload would be:

("LTPS", (MAP, (STRING, STRUCT, [
    ("eu-west", OrderedDict([("PSA", (STRING, "127.0.0.1")), ("PSP", (INT, 17502))])),
]))),
("LNP", (INT, 1)),     # drop from 10 -- one probe round, not ten

and it must ship together with an HTTPS listener on 127.0.0.1:17502 answering /qos/qos. The OSDK side degrades gracefully in the meantime: GetPingSiteAliasList impl 0x1472d17800x14728c4c0 explicitly returns an empty script array when the container is NULL or its count ([container+0x28]) is 0, so an empty ping-site list yields "" / [] rather than an error.

FIX 4 — instrumentation to buy, cheaply, what the next iteration will need

Independent of the above, and worth landing with FIX 1:

  • Sub-second timestamps in the responder log. Current second-resolution timestamps cannot resolve the RPC round trip, which is the whole question for a timeout bug. (Measured for this document: the pure-CPU cost of heat2.dump + hexdump on the 772-byte reply is only 0.08 ms/iter, and raw.sendall(out) at line 1252 already runs before the TX logging at 1254-1260 — so responder latency is almost certainly not a contributing factor. Timestamps would confirm that outright.)
  • Log the redirector request body. blaze_responder_v3b.py:1338 logs only the request line and build_redirect_response() (~:1274) ignores the request entirely, so the responder cannot actually substantiate "session 2 asked for the SAME service" — that was established from the byte-identical preAuth requests, not from the redirector.
  • Log FIN vs RST and bytes-read-before-close on the Blaze socket. The recv-loop EOF at blaze_responder_v3b.py:1195-1270 currently reports a bare closed for every case.
  • Timestamp /tmp/lsx.log. The four GetAuthCode issuances (ids 25-28) cannot currently be aligned with the Blaze log.

3. What still needs a live experiment

Static RE has taken this as far as it can go. Everything below needs the game.

3.1 The decisive one — apply FIX 1, then re-enter online (no relaunch strictly required, but cleaner)

  1. Edit the three CONF lines, restart the responder only.
  2. Relaunch FIFA17 (cleanest: guarantees a fresh ConnectionManager with CM+0x278 = 10000 from the ctor and CM+0xb3c = 0).
  3. Immediately after the boot login, read live CM+0xd1c / CM+0x278 / CM+0xd28. Gate: they must read 20000 / 30000 / 90000. If they do not, the fix did not apply and nothing else matters.
  4. Then have the user go online. Gate: the log must show Util::ping after our preAuth reply on that connection, then Authentication::login.

Two clean outcomes, both informative:

  • Connection survives → FIX 1 was the whole bug. Proceed to whatever the next gate is (likely the constant AuthCode OPENFUT-0000… we return on every LSX GetAuthCode).
  • Connection still dies → the zero timeouts were real but not sufficient. Go to §3.2.

3.2 If it still dies — read the error code the SDK hands the callback

This is the single highest-value remaining measurement and it is small. Breakpoint (or single-step-and-read) at 0x146e1cf10 (onPreAuthResponse entry) and capture r8d, or at the branch 0x146e1cf33 (test r8d,r8d / je 0x146e1cf82). That one integer separates:

  • callback entered with a non-zero errorCode → the SDK generated the failure itself (timeout, transport) — chase the failure dispatcher 0x146e18170 (callers 0x146e1990c, 0x146e1c1c2, 0x146e1cf77, 0x146e1d2de) and the connect-failure handler 0x146e1cc70;
  • callback never entered at all → the socket was torn down before dispatch, i.e. exactly the timeout-expiry story, and the remaining question is who called disconnect.

Pair it with a live read of LoginStateLogin (0x43d189d8, vtable 0x14395c180) at the moment of failure: if +0x80 becomes 0x14395c6c8 the state-machine selector fired and +0x260 names the outstanding async op (states {3,4,6,10,12,13,14,15,23,25} map to OSDK_LOST_CON_TO_EA); if +0x80 stays 0x143951fb0 with +0x7c == 0, it was the transport-disconnect popup at 0x1471aa220.

3.3 Open items that only dynamic work can close

# Question Why it is open How to close it
1 Who reads CM+0xd28 and CM+0x278? A disp32 scan for 0x00000d28 / 0x00000278 across 0x144ed3000-0x14a000000 finds only the write sites (0x146e1d163, 0x146e1d12b). Readers are inside VM-mutated functions. So "0 ⇒ close immediately" is inferred from behaviour (all three connections die at first idle), not proven by disassembly. FIX 1 is itself the experiment. If the connections stop dying, the inference was right.
2 What writes CM+0xb3c (the preAuth-skip flag)? Byte-pattern scans for c6 8x 3c 0b 00 00 over 0x144ed3000-0x149000000 find no writer; only the read at 0x146e1cbed. Set from a virtualized function. Watchpoint on CM+0xb3c. Knowing it would let us distinguish reconnect from fresh connect — and possibly force every connection onto the working preAuth-skipping path.
3 Two functions on the QoS completion path are VM stubs: 0x146e19b30 (push rcx; lea rcx,[0x148775779]; jmp 0x14e02d517) and the QoS-retrieved callback 0x146e1c5e0 (push rcx; lea rcx,[0x146fdea3d]; jmp 0x149875d6c). Same for the writer of CM+0x1f60 (only the read at 0x146e1c3c2 exists). Denuvo/VMProtect-style. Not statically decompilable. Session 1 provably ran both to completion, so they are not fatal. Only worth tracing if FIX 3 ever becomes necessary.
4 Does the client actually probe :17502 once LTPS is populated? Never tested — LTPS has always been empty; nothing has ever opened that port (/proc/net/* clean). tcpdump -i lo port 17502 during the first attempt with a populated LTPS. Only relevant under FIX 3.
5 What is disconnect reason == 2? The dispatcher 0x14718dda0 tail-jumps to 0x1471aa220 passing its own arg; reasons 3/4/5 additionally emit a NETW/LoginError telemetry event (string 0x14395c840). The producer of the enum was not walked. Static chase not completed. Backtrace from 0x14718dda0 at the moment of failure. Almost certainly "connection to Blaze lost".
6 Blaze error codes 0x000B0001 and 0x000C0001 are the only two special-cased in OnBlazeError 0x1471d53e0 (cmp edx,0xb0001; je / cmp edx,0xc0001; je) — i.e. the only two failures FIFA treats as recoverable. Not decoded. Would tell us which failures we can safely provoke. Component 0x000b is live in our traffic (RX #4001 cmd:0x0a28). Decode against the component table; low priority.
7 The constant AuthCode. /tmp/lsx.log shows four GetAuthCode issuances (ids 25-28) and we return the same OPENFUT-0000… every time. Untested — we have never survived past preAuth on the online connection. This is the most likely next gate once FIX 1 lands. Have it in mind, do not pre-emptively change it.

4. One-line summary for the commit message

blaze: send CONF durations as unit-suffixed strings ("30s"), not bare microseconds

FIFA17's TimeValue getter (ConnMgr vt+0x58, 0x146e1bda0) discards the parse
result from 0x1479b2d50, which rejects unit-less integers.  pingPeriod,
defaultRequestTimeout and connIdleTimeout were all landing as 0 (live:
CM+0xd1c=15000 via clamp, CM+0x278=0, CM+0xd28=0), zeroing the request/idle
timeouts for the whole ConnectionManager lifetime and tearing down every
connection at its first idle moment -- including the go-online preAuth, whose
response callback never ran (proved by CM+0x11a4 still holding session 1's
16:38:06 tick).  Not a QoS/ping-site issue: LTPS-empty is a benign first-class
branch and nothing ever contacted :17502.