fifa17-recon: package the working offline FUT backend
Emulates FIFA 17's full online + Ultimate Team stack against an offline,
clean-room backend (no EA servers). Proven end-to-end 2026-08-01:
Origin login -> Blaze login -> device-trust -> the FUT hub.
Package:
- tools/openfut-fut.sh one-command orchestrator (start/stop/status/restart)
- tools/root_arm.sh idempotent host arm (sysctls, DNAT, /etc/hosts easw)
- tools/{lsx_responder_v2,blaze_responder_v3b,roster_server,utas_server,autopatch}.py
the 5 servers (Origin LSX :4216, Blaze :42127/42130/42131, roster :8081,
FUT/UTAS :8099) + heat2.py (Fire2/Heat2 TDF codec)
- FUT-RUNBOOK.md runbook + gate-ladder troubleshooting
- docs/, tools/login_dump/*.md the reverse-engineering write-ups
All findings are clean-room, from binaries we own; nothing from any leak.
The wire protocol maps 1:1 to FIFA 23.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PN5bmpDVQR1aXgefyWAt7o
This commit is contained in:
@@ -0,0 +1,496 @@
|
||||
# QOS_CONNECT_PLAN — "Unable to connect to the EA servers"
|
||||
|
||||
**Date:** 2026-07-31 · **Target:** FIFA17.exe (decrypted, VMA==runtime VA), live pid 50676
|
||||
**Sources:** decrypted image + `/proc/50676/mem` + `/tmp/blaze_responder.log` + `/tmp/lsx.log` +
|
||||
`blaze_responder_v3b.py`. **Clean-room: no leaked EA source was consulted.**
|
||||
|
||||
---
|
||||
|
||||
## 0. Verdict up front
|
||||
|
||||
> **The QoS hypothesis in BRIEF5 is WRONG. The empty `LTPS` and the unserved `127.0.0.1:17502` are
|
||||
> both harmless and were never even reached. The bug is a *silent type error in our `CONF` map*:
|
||||
> three of the values we send are TimeValue-typed and FIFA's duration parser rejects bare integers,
|
||||
> so `pingPeriod`, `defaultRequestTimeout` and `connIdleTimeout` all land as `0`.
|
||||
> The one-line class of fix is `"30000000"` → `"30s"`.**
|
||||
|
||||
BRIEF5's premise is also factually wrong and should be retired: **session 1 did not stay up for
|
||||
~2.5 minutes of pings.** It died at `16:38:07`, one second after the login burst, and the whole log
|
||||
contains exactly one `Util::ping` ever. The "2.5 minutes" is dead air between two failed connections.
|
||||
|
||||
Three of the four independent reverser passes converged on this cause. Everything load-bearing below
|
||||
was **re-verified from scratch** for this synthesis (live reads + raw bytes re-disassembled), and two
|
||||
new pieces of proof were added that no individual report had — see §1.5 and §1.6.
|
||||
|
||||
---
|
||||
|
||||
## 1. THE CAUSE — end to end
|
||||
|
||||
### 1.1 What we actually send
|
||||
|
||||
`blaze_responder_v3b.py:400-406`, inside `PreAuthResponse.CONF` (a `map<string,string>`):
|
||||
|
||||
```python
|
||||
("autoReconnectEnabled", "1"), # OK — read via atoi
|
||||
("connIdleTimeout", "90000000"), # BROKEN — TimeValue
|
||||
("defaultRequestTimeout", "30000000"), # BROKEN — TimeValue
|
||||
("maxReconnectAttempts", "5"), # OK — read via atoi
|
||||
("pingPeriod", "20000000"), # BROKEN — TimeValue
|
||||
```
|
||||
|
||||
We wrote microseconds as bare integers. FIFA reads these five keys through **two different vtable
|
||||
getters**, and only one of them accepts a bare integer.
|
||||
|
||||
### 1.2 The two getters
|
||||
|
||||
`ConnectionManager` vtable (live vptr `0x1438a0850`):
|
||||
|
||||
| Slot | Function | Semantics |
|
||||
|---|---|---|
|
||||
| `vt+0x48` | `0x146e1bb80`-adjacent | raw `getConfigString` |
|
||||
| `vt+0x50` | `0x146e1bb80` | **uint32** — string → `atoi` (import `0x148e21ad8`). **Works.** |
|
||||
| `vt+0x58` | `0x146e1bda0` | **TimeValue** — string → duration parser `0x1479b2d50`. **Broken for us.** |
|
||||
|
||||
`0x146e1bda0` re-disassembled live (raw bytes read out of pid 50676 for this document):
|
||||
|
||||
```
|
||||
146e1bda0: 53 push rbx
|
||||
146e1bda1: 4883ec20 sub rsp,0x20
|
||||
146e1bda5: 488b01 mov rax,[rcx]
|
||||
146e1bda8: 4c89c3 mov rbx,r8
|
||||
146e1bdab: 4c8d442430 lea r8,[rsp+0x30]
|
||||
146e1bdb0: ff5048 call [rax+0x48] ; getConfigString
|
||||
146e1bdb3: 84c0 test al,al
|
||||
146e1bdb5: 7506 jne 0x146e1bdbd
|
||||
146e1bdb7: 4883c420 5b c3 add rsp,0x20; pop rbx; ret ; key absent -> false
|
||||
146e1bdbd: 488b542430 mov rdx,[rsp+0x30]
|
||||
146e1bdc2: 4889d9 mov rcx,rbx
|
||||
146e1bdc5: e8866fb900 call 0x1479b2d50 ; TimeValue::parse <-- return value
|
||||
146e1bdca: b001 mov al,0x1 ; <-- CLOBBERED, UNCONDITIONAL "success"
|
||||
146e1bdcc: 4883c420 5b c3 add rsp,0x20; pop rbx; ret
|
||||
```
|
||||
|
||||
Line `146e1bdca: b0 01` is the whole problem: **the parser's boolean result is thrown away.** The
|
||||
caller is told the value was applied. It was not.
|
||||
|
||||
### 1.3 The parser rejects us
|
||||
|
||||
`0x1479b2d50` accumulates digits, then dispatches on the byte that terminates the run. Raw bytes at
|
||||
`0x1479b2db4`, read live:
|
||||
|
||||
```
|
||||
1479b2db4: 80f964 7435 cmp cl,'d' ; je
|
||||
1479b2db9: 80f968 742b cmp cl,'h' ; je
|
||||
1479b2dbe: 80f96d 7414 cmp cl,'m' ; je (+ 's' lookahead -> ms)
|
||||
1479b2dc3: 80f973 740a cmp cl,'s' ; je
|
||||
1479b2dc8: 80f979 7535 cmp cl,'y' ; jne 0x1479b2e02
|
||||
...
|
||||
1479b2e02: 30c0 xor al,al ; return false
|
||||
1479b2e04: eb50 jmp 0x1479b2e56 ; SKIPS the store at 0x1479b2e46
|
||||
```
|
||||
|
||||
Accepted grammar: optional `-`, then `:`-separated `<digits><unit>` with unit ∈ `{y, d, h, m, ms, s}`.
|
||||
A **NUL terminator with no unit** — i.e. exactly `"20000000"` — falls straight to `0x1479b2e02`,
|
||||
returns false, and **never executes `mov [r14],rax`**. The caller pre-zeroes each out-slot
|
||||
(`0x146e1d0af`, `0x146e1d10c`, `0x146e1d143`, all `mov QWORD PTR [rbp+N],rdi` with `rdi=0`), so the
|
||||
value the SDK ends up applying is **0**.
|
||||
|
||||
### 1.4 Live proof that all three landed as zero
|
||||
|
||||
Read out of pid 50676 for this document. `ConnectionManager = 0x43c47ff0` (vptr `0x1438a0850` ✓,
|
||||
`CM+0x1288 = "Blaze 15.1.1.3.0 (OpenFUT)\n"` — our own `SVER`, so this is the object our reply
|
||||
configured):
|
||||
|
||||
| Field | Offset | Written at | **Live value** | What we intended |
|
||||
|---|---|---|---|---|
|
||||
| `pingPeriod` | `CM+0xd1c` | `0x146e1d0c9` / `0x146e1d0f5` | **15000** | 20000 |
|
||||
| `defaultRequestTimeout` | `CM+0x278` | `0x146e1d12c` | **0** | 30000 |
|
||||
| `connIdleTimeout` | `CM+0xd28` | `0x146e1d164` | **0** | 90000 |
|
||||
| `maxReconnectAttempts` | `CM+0xc5c` | `0x146e1d1b6` | 5 ✓ | 5 |
|
||||
| `autoReconnectEnabled` | `CM+0x11b7` | `0x146e1d188` | 1 ✓ | 1 |
|
||||
|
||||
The split is exactly along the getter boundary: **every `atoi` key took, every TimeValue key was
|
||||
dropped.** That is not a coincidence, it is the signature.
|
||||
|
||||
The `15000` is *not* a "key missing" default — it is the clamp. Bytes at `0x146e1d0a5`, live:
|
||||
|
||||
```
|
||||
146e1d0a5: 488d1584 39a8fc lea rdx,[rip-0x357867c] ; -> 0x1438a0a30 = "pingPeriod"
|
||||
146e1d0ac: 4889d9 mov rcx,rbx
|
||||
146e1d0af: 48897d07 mov [rbp+7],rdi ; pre-zero out-slot
|
||||
146e1d0b3: ff5058 call [rax+0x58] ; getConfigTimeValue -> always true
|
||||
146e1d0b6: 48be cff753e3a59bc420 movabs rsi,0x20c49ba5e353f7cf ; /1000 magic
|
||||
146e1d0c0: 84c0 test al,al
|
||||
146e1d0c2: 750d jne 0x146e1d0d1
|
||||
146e1d0c4: b8983a0000 mov eax,0x3a98 ; 15000 = key-absent default
|
||||
146e1d0c9: 89831c0d0000 mov [rbx+0xd1c],eax
|
||||
```
|
||||
|
||||
`al` is always 1, so we take the `jne` — the parsed value (0) is divided by 1000 → 0, then the
|
||||
`cmovl` at `0x146e1d0f2` clamps anything below 1000 up to `0x3a98`. Reaching 15000 *through the
|
||||
clamp* is positive proof the parse produced 0. The `lea` also confirms the key name byte-for-byte:
|
||||
`0x1438a0a30` is exactly the address our own source comment on line 406 cites.
|
||||
|
||||
### 1.5 NEW PROOF — session 2's response callback provably never ran
|
||||
|
||||
`onPreAuthResponse` (`0x146e1cf10`) has **no branch on its success path that can skip
|
||||
`0x146e1d1cb: call 0x146e1e460` (sendPing)**. So "did a ping happen" is a perfect oracle for "did the
|
||||
callback run". The log shows exactly one ping ever (`RX #3984`, 16:38:06, conn 43427) and none in
|
||||
session 2.
|
||||
|
||||
`onPingResponse` (`0x146e1d290`) stamps `CM+0x11a0 = response STIM` and `CM+0x11a4 = local tick`.
|
||||
Live values, decoded for this document:
|
||||
|
||||
- `CM+0x11a0 = 1785541086` → **`2026-07-31 16:38:06`**, and our responder builds `STIM` as
|
||||
`int(time.time())` (`blaze_responder_v3b.py:564`) — this is literally the timestamp *we* generated
|
||||
and sent in session 1.
|
||||
- `CM+0x11a4 = 14437471 ms`. This is a system-monotonic tick, not process-relative. Against
|
||||
`/proc/uptime = 16638.9 s` at wallclock `17:14:48`, it back-solves to **`16:38:06`**.
|
||||
|
||||
Two independent clocks, both landing on **16:38:06**, not 16:40:34. The preAuth response callback
|
||||
last completed in session 1 and **did not run at all in session 2** — even though our reply was
|
||||
byte-identical (`cmp -l rx_3983… rx_4009…` differs in one byte, offset 13 = msgNum) and was fully
|
||||
received (the raw 772-byte payload is still sitting in the recv buffer at `0x43c4e2xx-0x43c4e373`).
|
||||
|
||||
**The teardown therefore precedes the dispatch of our reply. Nothing about the reply's *content* is
|
||||
being rejected.** That kills the "QoS field validation" theory outright — the QoS fields are parsed
|
||||
*inside* the callback that never ran.
|
||||
|
||||
### 1.6 NEW PROOF — why session 1 survived preAuth and session 2 did not
|
||||
|
||||
This is the asymmetry BRIEF5 asked about, and it falls straight out of §1.4.
|
||||
|
||||
`ConnectionManager::ctor` at `0x146e187ea`:
|
||||
|
||||
```
|
||||
146e187ea: 488b5608 mov rdx,[rsi+0x8] ; BlazeHub
|
||||
146e187ee: 8b8244050000 mov eax,[rdx+0x544]
|
||||
146e187f4: 89867802 0000 mov [rsi+0x278],eax ; defaultRequestTimeout = hub+0x544
|
||||
```
|
||||
|
||||
Live `hub = 0x43c47330`, `hub+0x544 = 10000`. So a freshly-constructed ConnectionManager has a
|
||||
**healthy 10000 ms request timeout** — and `CM+0x278` is only ever overwritten **inside the preAuth
|
||||
*response* callback** (`0x146e1d12c`).
|
||||
|
||||
| | `CM+0x278` while its own preAuth RPC was outstanding | Outcome |
|
||||
|---|---|---|
|
||||
| conn A (43427, 16:38:06) | **10000 ms** (ctor default, not yet clobbered) | preAuth callback ran → ping sent → login burst |
|
||||
| conn C (35037, 16:40:34) | **0** (clobbered by conn A's callback, persists for the CM's lifetime) | preAuth job expires before the reply is dispatched → teardown |
|
||||
|
||||
**Our own preAuth reply poisons the timeout that the *next* preAuth needs.** Session 1 worked *because
|
||||
it ran before our config was applied*. That is why the identical bytes produce opposite outcomes, and
|
||||
why no amount of varying the QOSS content would ever have helped.
|
||||
|
||||
The same zero explains the rest of the log: with `connIdleTimeout = 0`, **every** connection dies at
|
||||
its first idle moment — conn A right after the ping reply, conn B (`43131`) right after `TX #4008.0`,
|
||||
the last frame of the login burst, conn C right after our preAuth reply. Conn B's 24 RPCs all
|
||||
succeeded because localhost replies land inside the same SDK tick they were issued on, so the expiry
|
||||
sweep never saw them outstanding.
|
||||
|
||||
Conn B also explains itself: `onSocketConnected` (`0x146e1cb60`) gates preAuth on a flag —
|
||||
|
||||
```
|
||||
146e1cbed: 4038bb3c0b0000 cmp BYTE PTR [rbx+0xb3c],dil ; dil = 0 (errorCode)
|
||||
146e1cbf4: 7507 jne 0x146e1cbfd
|
||||
146e1cbf6: e8f2140000 call 0x146e1e0f0 ; sendPreAuthRequest
|
||||
146e1cbfb: eb0c jmp 0x146e1cc09
|
||||
146e1cbfd: c68380110000 01 mov BYTE PTR [rbx+0x1180],1
|
||||
146e1cc04: e854180000 call 0x146e1e460 ; sendPing only
|
||||
```
|
||||
|
||||
Conn B skipped preAuth (its first frame is `fetchClientConfig msgNum=2`) → `CM+0xb3c` was non-zero →
|
||||
conn B was an **SDK auto-reconnect** (`autoReconnectEnabled=1`, `maxReconnectAttempts=5`, both live
|
||||
✓). `CM+0xb3c` reads **0** now, so the flag is cleared on full teardown — consistent with conn C
|
||||
re-sending preAuth. And `msgNum` is one monotonic counter `0…26` across all three sockets, with live
|
||||
`CM+0xc50 = 26` ✓ — **there is only one ConnectionManager**, so "session 2 is a different kind of
|
||||
connection" is false.
|
||||
|
||||
### 1.7 The QoS path is exonerated, with disasm
|
||||
|
||||
`QosManager::startLatencyProbes` `0x146e1dae0`, re-disassembled live:
|
||||
|
||||
```
|
||||
146e1dae0: 53 push rbx
|
||||
146e1dae1: 4883ec20 sub rsp,0x20
|
||||
146e1dae5: 488b8120020000 mov rax,[rcx+0x220] ; LTPS end
|
||||
146e1daec: 4889cb mov rbx,rcx
|
||||
146e1daef: 48398118020000 cmp [rcx+0x218],rax ; LTPS begin
|
||||
146e1daf6: 7508 jne 0x146e1db00
|
||||
146e1daf8: 30c0 xor al,al ; EMPTY -> false, benign, no side effects
|
||||
146e1dafa: 4883c420 5b c3 add rsp,0x20; pop rbx; ret
|
||||
```
|
||||
|
||||
Its caller `QosManager::initialize` `0x146e1c3f0` has a dedicated no-probes-started branch:
|
||||
|
||||
```
|
||||
146e1c5aa: e831150000 call 0x146e1dae0
|
||||
146e1c5af: 84c0 test al,al
|
||||
146e1c5b1: 7508 jne 0x146e1c5bb
|
||||
146e1c5b3: 4889f9 mov rcx,rdi
|
||||
146e1c5b6: e875d5ffff call 0x146e19b30 ; normal QoS-completion fn
|
||||
146e1c5bb: c6471001 mov BYTE PTR [rdi+0x10],1 ; initialized = true
|
||||
```
|
||||
|
||||
Live `QosManager = CM+0x1db0 = 0x43c49da0` (vptr `0x1438a07c8` ✓) proves that exact path ran to
|
||||
completion:
|
||||
|
||||
| Field | Live | Meaning |
|
||||
|---|---|---|
|
||||
| `Q+0x10` | **1** | `initialize()` reached its final store — completed |
|
||||
| `Q+0x20` | 0 | both QoS tests off — our `enableQos*Test=false` parsed fine |
|
||||
| `Q+0x218` / `Q+0x220` | 0 / 0 | LTPS empty → probe loop `0x146e1c4bb..0x146e1c4f0` never ran |
|
||||
| `Q+0x18` | 0 | no DirtySDK `QosApi` object exists |
|
||||
| `Q+0x1c8` / `Q+0x1f8` / `Q+0x250` | 10 / 17502 / 5000000 | our `LNP` / `BWPS.PSP` / `TIME` arrived intact |
|
||||
| `hub+0x53f` | 1 | QoS is *enabled*; the early-out at `0x146e1c44b` was **not** taken |
|
||||
| `CM+0x1f60` | 1 | SDK's own "QoS ready" flag is set |
|
||||
|
||||
So QoS ran, with QoS enabled, with an empty ping-site list, and **reported ready**.
|
||||
|
||||
**Nothing ever contacted 17502.** Verified for this document: no `445e` in `/proc/net/{tcp,udp,tcp6,udp6}`,
|
||||
and `ss -lntup` shows our responder listening only on 42127 / 42130 / 42131. The `BWPS` port is merely
|
||||
the `QosApi` *create* argument (`0x146e1975a: movzx r8d,WORD PTR [rbx+0x1f8]`), and the real probe
|
||||
transport is HTTPS (`https://%s:%u/qos/qos` @ `0x143aa75c0`) driven **only** by `LTPS` entries, of
|
||||
which there are none.
|
||||
|
||||
`PinQosError_ReferenceEvent` @`0x1439de5c8` has zero code xrefs (it sits in a telemetry name table
|
||||
next to `CM_WonTournament`, `PS3TOSAccepted`, …) and `POW:sConnectionManager` @`0x143995b88` is an
|
||||
allocation tag in the **powdll loader** string block. Both anchors in BRIEF5 are red herrings.
|
||||
|
||||
### 1.8 The on-screen message corroborates all of the above
|
||||
|
||||
The popup text is loc key `OSDK_LOST_CON_TO_EA` @`0x14395c6c8` (resolved English on the heap at
|
||||
`0x7b438e0`). Selection site `0x1471aa220`, bytes re-read live at `0x1471aa286`:
|
||||
|
||||
```
|
||||
1471aa286: 83ba0c02000001 cmp DWORD PTR [rdx+0x20c],1
|
||||
1471aa28d: 488d0534247bfc lea rax,[rip-0x384dbcc] ; 0x14395c6c8 OSDK_LOST_CON_TO_EA
|
||||
1471aa294: 4c8d1d157d7afc lea r11,[rip-0x38582eb] ; 0x143951fb0 OSDK_A_R30
|
||||
1471aa29b: 4c0f44d8 cmove r11,rax
|
||||
```
|
||||
|
||||
Caller `0x1471a1dda` passes `rdx = cnncMgr+0x4c4`, so the tested field is `cnncMgr+0x6d0` — the
|
||||
**"an EA request is currently in flight" busy flag** (set to 1 immediately before each async Blaze op
|
||||
at `0x1471b50bc`, `…5115`, `…5202`, `…5282`, `…5322`, `…5cb1`, `0x1471b6578`; cleared on completion at
|
||||
`0x14717d7bb`, `0x147190d06`, `0x1471b4fa4`, `0x1471b501e`, `0x1471b54a8`, `0x1471b5c37`, `0x1471b65c6`,
|
||||
`0x1471d5486`, `0x1471d6919`).
|
||||
|
||||
So the message means, literally: **"the socket died while an RPC was outstanding."** Not "QoS
|
||||
validation failed", and not "login timed out" — the 30 s login timeout path
|
||||
(`LoginStateVersionCheck::Tick` `0x1471b4ee0`, `lea eax,[rdx+0x7530]`) resolves to table index 0 and
|
||||
yields `OSDK_A_R30`, a *different* on-screen string. The alternative state-machine selector also did
|
||||
not fire: the live `LoginStateLogin` object (`0x43d189d8`, vtable `0x14395c180`) still has
|
||||
`+0x7c errCode = 0` and `+0x80 = 0x143951fb0` (the ctor default), i.e. no error was ever latched.
|
||||
|
||||
The popup is therefore a pure transport-disconnect notice, and it is *exactly* what a request that
|
||||
expires under a 0 ms timeout would produce.
|
||||
|
||||
### 1.9 Causal chain, one paragraph
|
||||
|
||||
We send `defaultRequestTimeout="30000000"` in `PreAuthResponse.CONF`. FIFA reads it with the TimeValue
|
||||
getter `0x146e1bda0`, whose parser `0x1479b2d50` requires a unit suffix and returns false on a bare
|
||||
integer; the getter discards that false (`146e1bdca: b0 01`) and the SDK applies the pre-zeroed
|
||||
out-slot, i.e. **0**, to `CM+0x278` — overwriting the ctor's healthy 10000 ms — plus `CM+0xd28 = 0`
|
||||
(idle timeout) and `CM+0xd1c` clamped to 15000. This happens *inside session 1's preAuth response
|
||||
callback*, so session 1's own preAuth was still protected by 10000 ms and completed normally. From
|
||||
that instant on, every connection on this ConnectionManager dies at its first idle moment, and any RPC
|
||||
that is not answered within the issuing tick is expired immediately. When the user goes online at
|
||||
16:40:34, conn C issues `Util::preAuth` under a 0 ms timeout; the job is expired before our reply is
|
||||
dispatched, the callback never runs (proved by the 16:38:06 ping tick), the SDK tears the socket down
|
||||
with `cnncMgr+0x6d0 == 1`, and `0x1471aa220` selects `OSDK_LOST_CON_TO_EA`: **"Unable to connect to
|
||||
the EA servers at this time."**
|
||||
|
||||
---
|
||||
|
||||
## 2. THE FIX — ranked and minimal
|
||||
|
||||
### FIX 1 — REQUIRED, proven, 3 lines, no new listener or thread
|
||||
|
||||
`blaze_responder_v3b.py:401-406`. Change bare microsecond integers to unit-suffixed duration strings:
|
||||
|
||||
```python
|
||||
("connIdleTimeout", "90s"), # was "90000000" -> parsed to 0
|
||||
("defaultRequestTimeout", "30s"), # was "30000000" -> parsed to 0
|
||||
("pingPeriod", "20s"), # was "20000000" -> clamped to 15000
|
||||
```
|
||||
|
||||
`"90000ms"` / `"30000ms"` / `"20000ms"` are equally valid. Leave `maxReconnectAttempts` and
|
||||
`autoReconnectEnabled` as plain integers — they go through the `atoi` getter and already work.
|
||||
|
||||
Grammar accepted by `0x1479b2d50`: optional leading `-`, then `:`-separated `<digits><unit>` with
|
||||
unit ∈ `{y, d, h, m, ms, s}`; combined as `((((y*365+d)*24+h)*60+m)*60+s)*1000+ms`, then ×1000 → µs.
|
||||
Corroborating precedent from the shipped image: the default for `QosConfigInfo.timeout` is the literal
|
||||
string `"5s"` @ `~0x14388cc90`.
|
||||
|
||||
While editing, add a comment recording the getter split so this class of bug cannot recur:
|
||||
|
||||
```python
|
||||
# CONF value TYPES matter. Keys read via ConnMgr vt+0x50 (0x146e1bb80) go through
|
||||
# atoi -> plain integers are fine. Keys read via vt+0x58 (0x146e1bda0) go through
|
||||
# TimeValue::parse (0x1479b2d50), which REQUIRES a unit suffix (y/d/h/m/ms/s) and
|
||||
# silently yields 0 for a bare integer -- 0x146e1bdca discards the parser's bool.
|
||||
```
|
||||
|
||||
**Non-interactive verification (no relaunch needed to check the *parse*, only a reconnect):**
|
||||
read the live ConnectionManager and assert the three fields flipped.
|
||||
|
||||
```bash
|
||||
PID=$(pgrep -x FIFA17.exe)
|
||||
python3 - <<'EOF'
|
||||
import struct
|
||||
PID=<pid>
|
||||
f=open("/proc/%d/mem"%PID,"rb")
|
||||
def u32(va): f.seek(va); return struct.unpack("<I",f.read(4))[0]
|
||||
CM=0x43c47ff0 # re-locate: scan rw memory for qword == 0x1438a0850
|
||||
print("pingPeriod CM+0xd1c =", u32(CM+0xd1c), "(want 20000, currently 15000)")
|
||||
print("defReqTO CM+0x278 =", u32(CM+0x278), "(want 30000, currently 0)")
|
||||
print("connIdleTO CM+0xd28 =", u32(CM+0xd28), "(want 90000, currently 0)")
|
||||
EOF
|
||||
```
|
||||
|
||||
> Note `CM=0x43c47ff0` is this process instance's address. After a relaunch, re-locate it by scanning
|
||||
> rw regions for the 8-byte vptr `0x1438a0850` (exactly one hit), or for the string
|
||||
> `Blaze 15.1.1.3.0 (OpenFUT)` and subtracting `0x1288`.
|
||||
|
||||
**Observable success signal (the one that matters):** on the go-online connection, after our
|
||||
`Util::preAuth` reply the client emits **`Util::ping` (component `0x0009`, cmd `0x0002`)** instead of
|
||||
`closed`, and proceeds to **`Authentication::login` (`0x0001`/`0x000a`)** on that same socket. A
|
||||
secondary signal: the boot connection stops dying ~1 s after the login burst and starts emitting a
|
||||
ping every ~20 s.
|
||||
|
||||
### FIX 2 — cheap, low risk, not the cause: close the CIDS gap
|
||||
|
||||
`blaze_responder_v3b.py:521-525`. We advertise `CIDS = (1, 4, 5, 7, 9, 15, 25, 28, 30722)`, but the
|
||||
session-1 log shows the client talking to three components we never listed:
|
||||
|
||||
| Component | Seen in log | Frame |
|
||||
|---|---|---|
|
||||
| `0x000a` (10) | `RX #3999 cmd:0x0005`, `RX #4008 cmd:0x0002` | UserSessions-adjacent |
|
||||
| `0x000b` (11) | `RX #4001 cmd:0x0a28` | — |
|
||||
| `0x08c9` (2249) | `RX #4005 cmd:0x0001` | — |
|
||||
|
||||
Add `10, 11, 0x08c9` to `COMPONENT_IDS`. Not the cause (the client sent to them regardless and we
|
||||
replied), but it is free and it removes a known divergence from a real server. Do this **after**
|
||||
FIX 1 lands and only if FIX 1 alone does not clear the gate, so the two changes stay separable.
|
||||
|
||||
### FIX 3 — DO NOT DO YET. Populating `LTPS` (and serving :17502)
|
||||
|
||||
**Explicitly deferred, and it is the one item that would need a new listener and thread.**
|
||||
|
||||
Reasons to leave `qos_config()` (`blaze_responder_v3b.py:527-538`) exactly as it is:
|
||||
|
||||
1. The empty-LTPS path is a *first-class, benign* branch (`0x146e1daf8: xor al,al`) and live memory
|
||||
proves it completed with QoS enabled (`Q+0x10=1`, `hub+0x53f=1`, `CM+0x1f60=1`). It blocks nothing.
|
||||
2. The QoS manager is initialized at `0x146e1d214`, **after** the `sendPing` at `0x146e1d1cb`, inside
|
||||
a callback that session 2 provably never entered (§1.5). It is downstream of the failure.
|
||||
3. Populating `LTPS` **creates** work: the per-site loop `0x146e1c4bb..0x146e1c4f0` would then run and
|
||||
fire real HTTPS `GET https://127.0.0.1:17502/qos/qos` requests. Serving those needs a **new TLS
|
||||
listener + thread** in the responder, with a cert the client accepts. We would be trading a
|
||||
non-problem for a real one.
|
||||
4. The Heat2 framing for `map<string, STRUCT>` is flagged **UNVERIFIED** in `heat2.py:204` (and the
|
||||
`STRUCT` value path in `_enc_value` likewise). A malformed TDF here would fail the client's decode
|
||||
of the *whole* PreAuthResponse — a strictly worse failure than the current one.
|
||||
|
||||
Only revisit if, **after FIX 1**, the game reaches login but then stalls on a ping-site-dependent
|
||||
feature. In that case the anchors are already mapped: `QosConfigInfo` copied to `QosManager+0x1b8`;
|
||||
`+0x1c8`=LNP, `+0x1d0`=BWPS (`QosPingSiteInfo` vtable `0x143889ba0`), `+0x200`=LTPS map (vtable
|
||||
`0x143889c48`, begin `+0x218` / end `+0x220`, **element stride 0x20**, int32 latency at `element+0x18`
|
||||
— stride and latency offset independently confirmed by the OSDK helper `0x1471c0c90`, which loops
|
||||
`add rax,0x20` writing `mov DWORD PTR [rax+0x18],0x3e8`). The minimal payload would be:
|
||||
|
||||
```python
|
||||
("LTPS", (MAP, (STRING, STRUCT, [
|
||||
("eu-west", OrderedDict([("PSA", (STRING, "127.0.0.1")), ("PSP", (INT, 17502))])),
|
||||
]))),
|
||||
("LNP", (INT, 1)), # drop from 10 -- one probe round, not ten
|
||||
```
|
||||
|
||||
and it must ship **together with** an HTTPS listener on `127.0.0.1:17502` answering `/qos/qos`.
|
||||
The OSDK side degrades gracefully in the meantime: `GetPingSiteAliasList` impl `0x1472d1780` →
|
||||
`0x14728c4c0` explicitly returns an empty script array when the container is NULL or its count
|
||||
(`[container+0x28]`) is 0, so an empty ping-site list yields `""` / `[]` rather than an error.
|
||||
|
||||
### FIX 4 — instrumentation to buy, cheaply, what the next iteration will need
|
||||
|
||||
Independent of the above, and worth landing with FIX 1:
|
||||
|
||||
- **Sub-second timestamps** in the responder log. Current second-resolution timestamps cannot resolve
|
||||
the RPC round trip, which is the whole question for a timeout bug. (Measured for this document: the
|
||||
pure-CPU cost of `heat2.dump` + `hexdump` on the 772-byte reply is only **0.08 ms/iter**, and
|
||||
`raw.sendall(out)` at line 1252 already runs *before* the TX logging at 1254-1260 — so responder
|
||||
latency is almost certainly not a contributing factor. Timestamps would confirm that outright.)
|
||||
- **Log the redirector request body.** `blaze_responder_v3b.py:1338` logs only the request line and
|
||||
`build_redirect_response()` (~:1274) ignores the request entirely, so the responder cannot actually
|
||||
substantiate "session 2 asked for the SAME service" — that was established from the byte-identical
|
||||
preAuth requests, not from the redirector.
|
||||
- **Log FIN vs RST and bytes-read-before-close** on the Blaze socket. The recv-loop EOF at
|
||||
`blaze_responder_v3b.py:1195-1270` currently reports a bare `closed` for every case.
|
||||
- **Timestamp `/tmp/lsx.log`.** The four `GetAuthCode` issuances (ids 25-28) cannot currently be
|
||||
aligned with the Blaze log.
|
||||
|
||||
---
|
||||
|
||||
## 3. What still needs a live experiment
|
||||
|
||||
Static RE has taken this as far as it can go. Everything below needs the game.
|
||||
|
||||
### 3.1 The decisive one — apply FIX 1, then re-enter online (no relaunch strictly required, but cleaner)
|
||||
|
||||
1. Edit the three CONF lines, restart **the responder only**.
|
||||
2. Relaunch FIFA17 (cleanest: guarantees a fresh ConnectionManager with `CM+0x278 = 10000` from the
|
||||
ctor and `CM+0xb3c = 0`).
|
||||
3. Immediately after the boot login, read live `CM+0xd1c` / `CM+0x278` / `CM+0xd28`. **Gate:** they
|
||||
must read `20000 / 30000 / 90000`. If they do not, the fix did not apply and nothing else matters.
|
||||
4. Then have the user go online. **Gate:** the log must show `Util::ping` after our preAuth reply on
|
||||
that connection, then `Authentication::login`.
|
||||
|
||||
Two clean outcomes, both informative:
|
||||
- **Connection survives → FIX 1 was the whole bug.** Proceed to whatever the next gate is (likely
|
||||
the constant `AuthCode` `OPENFUT-0000…` we return on every LSX `GetAuthCode`).
|
||||
- **Connection still dies → the zero timeouts were real but not sufficient.** Go to §3.2.
|
||||
|
||||
### 3.2 If it still dies — read the error code the SDK hands the callback
|
||||
|
||||
This is the single highest-value remaining measurement and it is small. Breakpoint (or
|
||||
single-step-and-read) at `0x146e1cf10` (`onPreAuthResponse` entry) and capture **`r8d`**, or at the
|
||||
branch `0x146e1cf33` (`test r8d,r8d` / `je 0x146e1cf82`). That one integer separates:
|
||||
|
||||
- **callback entered with a non-zero errorCode** → the SDK generated the failure itself (timeout,
|
||||
transport) — chase the failure dispatcher `0x146e18170` (callers `0x146e1990c`, `0x146e1c1c2`,
|
||||
`0x146e1cf77`, `0x146e1d2de`) and the connect-failure handler `0x146e1cc70`;
|
||||
- **callback never entered at all** → the socket was torn down before dispatch, i.e. exactly the
|
||||
timeout-expiry story, and the remaining question is *who* called `disconnect`.
|
||||
|
||||
Pair it with a live read of `LoginStateLogin` (`0x43d189d8`, vtable `0x14395c180`) **at the moment of
|
||||
failure**: if `+0x80` becomes `0x14395c6c8` the state-machine selector fired and `+0x260` names the
|
||||
outstanding async op (states `{3,4,6,10,12,13,14,15,23,25}` map to `OSDK_LOST_CON_TO_EA`); if `+0x80`
|
||||
stays `0x143951fb0` with `+0x7c == 0`, it was the transport-disconnect popup at `0x1471aa220`.
|
||||
|
||||
### 3.3 Open items that only dynamic work can close
|
||||
|
||||
| # | Question | Why it is open | How to close it |
|
||||
|---|---|---|---|
|
||||
| 1 | **Who reads `CM+0xd28` and `CM+0x278`?** A disp32 scan for `0x00000d28` / `0x00000278` across `0x144ed3000-0x14a000000` finds **only the write sites** (`0x146e1d163`, `0x146e1d12b`). | Readers are inside VM-mutated functions. So "0 ⇒ close immediately" is inferred from behaviour (all three connections die at first idle), **not** proven by disassembly. | FIX 1 is itself the experiment. If the connections stop dying, the inference was right. |
|
||||
| 2 | **What writes `CM+0xb3c`** (the preAuth-skip flag)? Byte-pattern scans for `c6 8x 3c 0b 00 00` over `0x144ed3000-0x149000000` find no writer; only the read at `0x146e1cbed`. | Set from a virtualized function. | Watchpoint on `CM+0xb3c`. Knowing it would let us distinguish reconnect from fresh connect — and possibly force *every* connection onto the working preAuth-skipping path. |
|
||||
| 3 | **Two functions on the QoS completion path are VM stubs**: `0x146e19b30` (`push rcx; lea rcx,[0x148775779]; jmp 0x14e02d517`) and the QoS-retrieved callback `0x146e1c5e0` (`push rcx; lea rcx,[0x146fdea3d]; jmp 0x149875d6c`). Same for the writer of `CM+0x1f60` (only the read at `0x146e1c3c2` exists). | Denuvo/VMProtect-style. Not statically decompilable. | Session 1 provably ran both to completion, so they are not fatal. Only worth tracing if FIX 3 ever becomes necessary. |
|
||||
| 4 | **Does the client actually probe `:17502` once LTPS is populated?** | Never tested — LTPS has always been empty; nothing has ever opened that port (`/proc/net/*` clean). | `tcpdump -i lo port 17502` during the first attempt with a populated LTPS. Only relevant under FIX 3. |
|
||||
| 5 | **What is disconnect `reason == 2`?** The dispatcher `0x14718dda0` tail-jumps to `0x1471aa220` passing its own arg; reasons 3/4/5 additionally emit a `NETW`/`LoginError` telemetry event (string `0x14395c840`). The producer of the enum was not walked. | Static chase not completed. | Backtrace from `0x14718dda0` at the moment of failure. Almost certainly "connection to Blaze lost". |
|
||||
| 6 | **Blaze error codes `0x000B0001` and `0x000C0001`** are the only two special-cased in `OnBlazeError` `0x1471d53e0` (`cmp edx,0xb0001; je` / `cmp edx,0xc0001; je`) — i.e. the only two failures FIFA treats as recoverable. Not decoded. | Would tell us which failures we can safely provoke. Component `0x000b` is live in our traffic (`RX #4001 cmd:0x0a28`). | Decode against the component table; low priority. |
|
||||
| 7 | **The constant `AuthCode`.** `/tmp/lsx.log` shows four `GetAuthCode` issuances (ids 25-28) and we return the same `OPENFUT-0000…` every time. | Untested — we have never survived past preAuth on the online connection. | This is the most likely **next** gate once FIX 1 lands. Have it in mind, do not pre-emptively change it. |
|
||||
|
||||
---
|
||||
|
||||
## 4. One-line summary for the commit message
|
||||
|
||||
```
|
||||
blaze: send CONF durations as unit-suffixed strings ("30s"), not bare microseconds
|
||||
|
||||
FIFA17's TimeValue getter (ConnMgr vt+0x58, 0x146e1bda0) discards the parse
|
||||
result from 0x1479b2d50, which rejects unit-less integers. pingPeriod,
|
||||
defaultRequestTimeout and connIdleTimeout were all landing as 0 (live:
|
||||
CM+0xd1c=15000 via clamp, CM+0x278=0, CM+0xd28=0), zeroing the request/idle
|
||||
timeouts for the whole ConnectionManager lifetime and tearing down every
|
||||
connection at its first idle moment -- including the go-online preAuth, whose
|
||||
response callback never ran (proved by CM+0x11a4 still holding session 1's
|
||||
16:38:06 tick). Not a QoS/ping-site issue: LTPS-empty is a benign first-class
|
||||
branch and nothing ever contacted :17502.
|
||||
```
|
||||
Reference in New Issue
Block a user