# Changelog ## 1.9.2-patch1 **Server/CMS-only connection-lifecycle hardening for #148 — NO Android APK, players stay on their current builds.** This strictly HELPS and de-risks, but is **NOT a guaranteed #148 fix**: the MAXHUB client-side reconnect failure and the disconnect synchronizer (edge conntrack / reporting) are separate, unproven-here tracks that may still require a client update / a Bold Sophos-edge review — **do not consider #148 fully closed on this patch alone.** ### Fixed / hardened — connection lifecycle (#148) - **The flap-limiter no longer quarantines legitimate PAIRED devices on reconnect churn.** A paired + authenticated device reconnecting is exempt from the 30-min quarantine escalation (a brief soft cooldown at most), so a repeated edge/NAT flush behind one SNAT IP can no longer be amplified into a self-inflicted fleet-wide lockout. Unpaired/abusive flapping is still quarantined (the attacker / unprovisioned-hammering case is unchanged). - **Marking a device offline now also closes its socket**, so DB-offline can't diverge from socket-state into a silent half-open the client is never told about. - **Faster half-open detection:** ping interval 30s → 15s (the pong TIMEOUT is kept at 30s so decode-loaded TV WebKits aren't falsely dropped) → dead-peer detection 60s → 45s on BOTH the server AND the client (the client inherits these via the handshake — **no APK needed**). - **TCP SO_KEEPALIVE** on every connection so a half-open TCP can't persist indefinitely at the OS layer. Server/CMS version only; ships no APK (versionCode still increments so a future player build is OTA-recognized). Docker: `ghcr.io/screentinker/screentinker:1.9.2-patch1` (pre-release — `:latest` stays at 1.9.2). ## 1.9.2 **⚠ Major internal hardening release (the "#146" rewrite) — large blast radius.** 1.9.2 rewrites the connection / maintenance / OTA hot paths to kill an event-loop death spiral, plus adds usage-metering (billing) and web-player fixes. If you bisect a regression to the 1.9.x line, 1.9.2 is the big one. Core invariant introduced: **no synchronous op may block the event loop for more than ~50ms**, ever. Every new subsystem has an env kill-switch. ### Fixed — maintenance / prune (the death-spiral root cause) - **Non-blocking, chunked, per-device `device_status_log` prune.** The old whole-table `ROW_NUMBER` sort froze boot for 40–48s at ~1M rows → healthcheck fail → restart loop that wiped in-memory throttle state → the spiral. Prune is now per-device, indexed, batched with `setImmediate` yields (`lib/chunked-prune.js`), async, re-entrant, and band-gated on the interval run (the startup prune is intentionally un-gated so a bloated table self-heals on first boot without freezing it). All table-growth sweeps (status-log, play-logs, provisioning, telemetry, lag) route through the chunked helper. New index `idx_devices_provisioning`. **Measured worst-case event-loop gap under the storm harness: <300ms across 300k rows (was 40–48s).** ### Fixed — reconnect / flap - **Per-device flap-rate limiter** (`lib/flap-limiter.js`): a device reconnecting faster than `CONNECT_RATE_MAX` (20) per `CONNECT_RATE_WINDOW_MS` (5min) is refused at the register gate, keyed via a **SNAT-safe identity chain** (device_id → fingerprint → token → one bounded global anon bucket) — **never by IP** (the whole fleet egresses one IP). After repeated trips a hard flapper is **quarantined IN-MEMORY for 30min and auto-clears** — it is NOT a durable DB block. - **Operator block kill-switch:** `POST /api/devices/:id/{block,unblock}` + a dashboard button; the block check resolves the effective device_id via the identity chain so a device_id-less reconnect of a blocked device is still caught. Takes effect on next register, no restart. - Also folded in: false-offline fixes (live-socket liveness beats a lagged heartbeat clock; evicted-socket re-arm race) and per-connection fail-fast so one device's handler throw can never exit the process. ### Fixed — OTA (SNAT-safe) - `/api/update/check` early-returns before any filesystem call when there's no offer; APK metadata is cached. `/download/apk` gains a **band-aware** global concurrency + rate guard that sheds with **503 Retry-After only under elevated/critical loop-lag** — under normal band, downloads serve freely (a coordinated fleet rollout is never staggered when healthy). All limiting is global/aggregate — **no per-IP limiting** (SNAT). ### Added — telemetry / logging / observability - Batched `event_loop_lag` inserts (buffered, flushed every 10s) and coalesced high-frequency logging (one summarized line per key per 30s; band *changes* stay immediate). - **Throughput counters** (running total + last-completed-window) in the `/api/status` debug block so a flapper/flood shows on the server itself (`flap.refusedLastWindow` climbing while `band=normal` = the limiter absorbing it cheaply). The debug block is now **admin-toggleable** (Admin tab, persisted, no restart; default follows `STATUS_DEBUG_ENABLED`). - **`devices_connected`** on `/api/status` (always-on): the live WS-socket count from the heartbeat connection map (NOT the lagging `devices.status='online'` column). ### Added — billing (usage metering) - **Billable Screens** metering per the ByteTinker–Bold agreement — the contractual system-of-record. A durable daily rollup (`device_usage_daily`) is accumulated incrementally off the heartbeat tick from live presence (retention-independent), pruned chunked. Exposed on a **dedicated, admin-gated `GET /api/billing/usage`** route (NOT on `/api/status`; billing is revenue data). Readable via an **owner-minted, revocable `billing:read` scoped token** (`scripts/mint-billing-token.js`) that authorizes billing-read and nothing else, OR a platform-admin session. See [`docs/billing.md`](docs/billing.md). ### Fixed — web player - **"Unchanged" refresh no longer drops the video.** On a reconnect the server re-emits `device:paired` while content is already playing; the player showed the idle "Waiting for content…" overlay unconditionally (covering live video; audio kept playing underneath) and the following "Playlist unchanged" left it up. Idle now shows only when genuinely idle, and an unchanged refresh is a strict no-op that leaves playback exactly as-is. - Hardened `PlayerMediaHealth` call sites to guard by **method** (not object) so a stale-cached player module can't throw `shouldShowIdle is not a function` and abort a socket handler. ### Added — translations - Italian (`it`) locale updated (#145). ## 1.9.2-beta1 — unreleased ### Fixed — server resilience (#142) - **A single flapping device can no longer saturate the event loop.** A new load-aware, per-device reconnect throttle (`lib/reconnect-throttle.js`) gates genuine reconnects *before* the heavy register work (DB writes + playlist build). The verdict is per-device; global event-loop lag only multiplies an already-flagged device's backoff and never throttles a healthy one. Hard ceiling + cold-start warm-up so a full-fleet reconnect after a deploy is never throttled. - **`device_status_log` growth is bounded.** Added `idx_device_status_log_device_ts`, a global retention sweep (`pruneStatusLog`, `STATUS_LOG_RETENTION_DAYS` default 3) covering removed/idle devices and the `offline_timeout` path, and de-duplicated the table's `CREATE TABLE`. - **`content-ack` spam de-duplicated.** Repeated identical `(device_id, content_id, status)` reports are suppressed within `CONTENT_ACK_DEDUP_MS` (default 10s). - **Provisioning cleanup window corrected.** Unclaimed provisioning devices are now swept after 24h (the code used `365 * 86400` — a year — contradicting its own comment). ### Added — observability (#142) - **Event-loop lag telemetry** via `perf_hooks.monitorEventLoopDelay()`. Sampled to a bounded `event_loop_lag` table (indexed + pruned, `LAG_TELEMETRY_RETENTION_DAYS`) and surfaced on `/api/status` as `loop_lag` (mean/p50/p99/max + band). ### Maintenance - Operators whose `device_status_log` is already bloated from a pre-1.9.2 deployment should reclaim disk with a **one-time manual `VACUUM`** in a maintenance window; retention now bounds further growth. Auto-VACUUM is intentionally not enabled. See [`docs/maintenance-device-status-log.md`](docs/maintenance-device-status-log.md). ## 1.9.1-beta3 — unreleased ### Fixed — Tizen player - **#118 Sticky "Not authenticated" banner.** On TV sleep/wake the socket reconnects and a heartbeat could fire on the fresh, not-yet-registered socket; the server rejected it with `device:auth-error`, which the player showed as a *sticky* toast over still-playing content (and, worse, dropped its saved credentials and re-paired). Heartbeats are now gated on a per-connection `authenticated` flag (set only between `device:registered` and `disconnect`/`auth-error`), the heartbeat timer is stopped on `connect`/`disconnect`/ `auth-error`, the stale banner is cleared on `device:registered`, and the `auth-error` toast is non-sticky so any transient case self-clears. - **#119 `app_version` stuck at `1.0.0`.** The hardcoded constant made every Tizen device report `1.0.0` regardless of the installed `.wgt`. The version now resolves at runtime from `config.xml` via the Tizen application API, with a fallback constant that `build-wgt.sh` stamps from `config.xml`'s `version=""`. ### Added — Tizen player - **Video walls (`wall:sync`).** The Tizen player now supports wall membership: when the payload carries `wall_config`, a new `WallController` positions the stage (vw/vh) as this screen's slice of the wall and drives the single-zone player as leader or follower. The leader broadcasts `wall:sync` at 4Hz; followers align their index and keep their video locked to the leader's clock with a latency-compensated drift controller (hard-seek past 0.3s, gentle ±3% playbackRate nudge past 0.05s), and request an immediate position on (re)connect via `wall:sync-request`. Mirrors the web player (the Android player has no wall support). Per-tile `rotation` is not applied yet (web-player parity). Wall emits are gated on auth + connection so a pre-register tick can't trip `device:auth-error`. - **Multi-zone layouts (Android parity).** The Tizen player now renders assigned layouts, not just fullscreen single-zone. A new `ZoneRenderer` (ports the Android `ZoneManager`) positions zones by percent geometry with `z_index`/`fit_mode`/background, groups assignments by `zone_id` (unassigned content goes to the first zone), and rotates each zone independently with the same per-item schedule gating (#74/#75). `app.js` selects the renderer from `payload.layout`; single-zone playback is unchanged. (Video walls `wall:sync` are still Android-only.) - **#121 Remote commands.** Added a `device:command` handler (`refresh`, `launch`, `screen_on`, `screen_off`, plus honest no-op toasts for `update`/`reboot`/`shutdown`, which need B2B/MDM privileges a sideloaded app lacks). Removed the dead `device:reload` listener (the server never emitted it) in favour of `device:command` `refresh`. - **#120 Dashboard preview.** Added `device:screenshot-request` / `device:remote-start` / `device:remote-stop`. Images capture for real; `