screentinker/docs/146-hardening-plan.md
ScreenTinker 106eddee52 docs(#146): event-loop hardening plan — blast-radius audit + failure model + sequenced plan
Phase 0 deliverable for the beta7 (alpha-only) hardening pass. Enumerates every
synchronous unbounded op (maintenance sweeps, register hot path, content-ack, OTA
endpoints, loop-lag telemetry, log volume, un-evicted per-device Maps), classifies
each exposed/mitigated with worst-case blocking, confirms the two-mechanism spiral
(whole-table prune freeze <-> restart-loop throttle wipe), and sequences items A-E.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-06-30 20:44:46 -05:00

129 lines
8.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# #146 — Final event-loop hardening pass (beta7, ALPHA-ONLY)
**Goal / acceptance bar:** a server that *cannot* be frozen or driven into a restart
loop by any single misbehaving device, a bloated table, or a SNAT'd request flood.
**Core invariant:** no synchronous operation may block the event loop for more than
~50ms, ever, regardless of table size or request rate. The system self-recovers to
healthy *without a process restart*.
Branch `fix/146-hardening` off the beta6 tip (keeps the crash-isolation from beta6;
beta7 = beta6 + this). **Do not bump/tag/release/deploy** — Dan bumps + pulls to alpha.
---
## (b) Confirmed failure model (verified against current code)
Two interlocked mechanisms, not one:
1. **TRIGGER** — a device flapping on a multi-second cadence (connect → ~5 content
acks → disconnect → reconnect every ~35s). Each cycle is expensive (register +
`buildPlaylistPayload` + acks) *and* writes one `device_status_log` row/cycle,
bloating the table over the retention window. The existing burst throttle
(`reconnectBaseMax=5` / `reconnectWindowMs=10s`, deviceSocket.js:425) does **not**
catch it: ~23 connects/10s passes clean. → **Item B**.
2. **AMPLIFIER**`pruneStatusLog()` (db/database.js:763) is a **whole-table
`ROW_NUMBER() OVER (PARTITION BY device_id …)` sort**. On the bloated table
(1,116,544 rows in the incident) it blocks the loop **4048s synchronously**
(confirmed by `mean=NaNms p99=0ms` lag samples beside `[status-log] pruned` lines).
It runs on the heartbeat interval **and at startup** → a bloated table freezes boot
→ healthcheck fails → restart. → **Item A**.
3. **INTERLOCK** — the reconnect throttle is in-memory and resets every restart, then
re-enters a 30s warmup where throttling is relaxed. The prune-induced restart loop
wipes the throttle's history every cycle, so the flapper is never caught. **Ending
the restart loop (Item A) is what lets the flap limiter (Item B) gain traction.**
---
## (a) Blast-radius audit — every synchronous, potentially-unbounded op
Legend: **EXPOSED** = can block >50ms under load/growth; **MITIGATED** = already bounded.
### Maintenance sweeps
| Path | Shape | Worst-case block | Reachable in outage? | Class | Fix |
|---|---|---|---|---|---|
| `pruneStatusLog` (db/database.js:763) | whole-table `ROW_NUMBER` sort | **4048s @ 1.1M rows** (O(N log N)) | YES (interval + startup) | **EXPOSED — the amplifier** | A |
| `pruneLag` (loop-lag.js:92) | indexed range `DELETE … sampled_at < ?` | O(deleted); one statement can delete all-old if stalled | YES (interval + startup) | **EXPOSED** | E |
| event_loop_lag INSERT (loop-lag.js:71) | sync INSERT per sample | small, but per-tick on the loop; worsens under DB contention | YES | **EXPOSED (freq)** | E |
| `pruneTelemetry(id)` (database.js:773) | per-device `NOT IN (SELECT … LIMIT 6000)`, idx_telemetry_device | one device's excess in one DELETE; **runs per-heartbeat** (deviceSocket.js:626) | partial | **EXPOSED (freq)** | A (helper + throttle) |
| play_logs 90d prune (heartbeat.js:64) | indexed range DELETE | O(>90d rows) one statement | YES (interval) | **EXPOSED** | A |
| `pruneProvisioningDevices` (heartbeat.js:94) | DELETE devices + FK cascade | provision-flood → big cascade delete | YES (interval) | **EXPOSED** | A |
| `pruneScreenshots(id)` (database.js:785) | per-device keep-newest-1 | ~rows-per-device | partial | MITIGATED | — |
| team/workspace invite prunes (heartbeat.js:73,78) | expiry DELETE, small tables | negligible | YES | MITIGATED (route for uniformity) | A(opt) |
### Hot paths
| Path | Worst-case | Class | Fix |
|---|---|---|---|
| `device:register` (deviceSocket.js:282) — token validate, blocked SELECT, evict, DB writes, `buildPlaylistPayload` | full cost per flap; burst throttle misses 35s flappers | **EXPOSED** | B |
| content-ack (deviceSocket.js:647) | limiter sheds the loop cost, **but `state` Map has no eviction** (unbounded device keys) | **EXPOSED (unbounded map)** | E |
| `/api/update/check` (server.js:585) | `resolveApkPath` `existsSync`×2 **unconditional** (even on rate-backoff no-offer); `console.log` per check | **EXPOSED** | C |
| `/download/apk` (server.js:734) | `resolveApkPath` per req; **IP-keyed** log throttle (SNAT-collapsed → hid the flood); no concurrency/rate/band cap on 8.8MB `sendFile` | **EXPOSED** | C |
### Unbounded in-memory structures (invariant: every Map bounded)
`content-ack state`, `lastPlayLogAt`, `lastReclaimRejectLogAt`, writer `lastWritten`
have **no eviction** — grow with distinct device_ids (a SNAT flood minting provisioning
ids inflates them). `otaDownloadLoggedAt` is **IP-keyed** (wrong under SNAT). Bounded
already: reconnect-throttle (#146 sweep), ota-breaker (#144 sweep), `pendingOfflines`,
`evictedSockets`. → **Item E adds a shared bounded-map sweep; Item C replaces the IP map.**
### Logging
Band lines, per-request OTA lines, repetitive "Device reconnected" — synchronous
`console.log` at storm rate blocks the loop. → **Item E** coalescing buffer.
---
## Identity fallback chain (applies to Items B & D — HARD CONSTRAINT)
SNAT: the whole fleet egresses as one IP (10.10.10.1). **Never key on IP.** Resolve
identity as: `device_id``fingerprint` (map via `device_fingerprints`→device_id when
it resolves, else the raw fingerprint string) → `device_token` → a **single bounded
GLOBAL anon bucket**. `fingerprint` is present in the `device:register` payload
(deviceSocket.js:289). An unidentifiable client is still bucketed (global anon) so an
anonymous flood is capped, never unthrottled. Mirrors `lib/ota-breaker`'s
device_id-or-version fallback. Extracted as `lib/device-identity.js`.
---
## (c) Sequenced implementation plan — isolated commits
- **A. Non-blocking maintenance.** New `lib/chunked-prune.js` (`chunkedPrune`: bounded
`LIMIT` batch DELETE + `setImmediate` yield between batches, optional band-gate,
re-entrancy guard). Rewrite `pruneStatusLog` → per-device indexed prune (keep newest
`statusLogMaxRowsPerDevice` per device via idx_device_status_log_device_ts; delete
older + past retention), chunked, async, band-gated on the interval / un-gated at
startup. Route play_logs, provisioning-cascade, telemetry, event_loop_lag prune
through the helper. `STATUS_LOG_PRUNE_BATCH=2000`. Callers fire-and-forget, no
stacking. **Tests:** correctness, ≥300k non-blocking in batches, band-gate, re-entrancy.
- **B. Flap-rate limiter.** New `lib/flap-limiter.js` — per-identity long-window connect
frequency (`CONNECT_RATE_WINDOW_MS=300000`, `CONNECT_RATE_MAX=20`), keyed via the
fallback chain, bounded (sweep + global anon bucket). Checked at the register gate
before heavy work; over-limit → backoff notice + disconnect. Optional auto-quarantine
via Item D block after N refusals. **Tests:** 4s-flapper refused, normal never,
two device_ids independent, fingerprint-bucketed, anon-bucketed, never IP.
- **C. OTA under SNAT.** Early-return before any FS on no-offer verdict; cache APK
resolution (path/size/mtime at boot + interval refresh); `/download/apk` global
concurrency + global rate + band shed (503 Retry-After); replace IP log-throttle with
a bounded aggregate per-window counter. **Tests:** rate-backoff ⇒ 0 fs calls; download
sheds 429/503 past global cap + under critical; nothing keys on IP.
- **D. Operator block.** Move blocked check behind the identity fallback chain (catch a
device_id-less reconnect of a blocked device via fingerprint→device_id); dashboard
block/unblock toggle (DB write + UI + next-register effect, no restart); in-code outage
procedure. **Tests:** blocked rejected cheaply at handshake; device_id-less+mapped-
fingerprint still rejected; unblock next-register.
- **E. Log/write self-protection.** `lib/log-coalescer.js` (dedup+count, interval flush,
bounded); buffer+batch event_loop_lag inserts; route its prune through `chunkedPrune`;
add a shared bounded-map sweep for the un-evicted per-device Maps. **Tests:** N identical
lines ⇒ 1 summarized, buffer bounded.
- **STORM HARNESS:** flapper + pre-bloated status_log + OTA flood from one SNAT IP +
a sweep, all at once; assert loop-lag never enters a multi-second freeze and the
server stays responsive (no NaN-sample block).
## Before/after worst-case blocking (targets)
| Hot path | Before | After (target) |
|---|---|---|
| pruneStatusLog @1.1M | 4048s | <50ms/batch, many batches |
| play_logs / provisioning / lag prune | O(all-old) one stmt | <50ms/batch |
| /api/update/check (no-offer) | existsSync×2 + log | 0 FS, coalesced log |
| /download/apk flood | unbounded sendFile + hidden | global-capped, 503 shed, visible |
| event_loop_lag telemetry | sync INSERT/sample | batched flush |
| device:register (flapper) | full build every ~4s | refused at gate, cheap |