screentinker/CHANGELOG.md

265 lines
17 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Changelog
## 1.9.2-patch2
**Server/CMS-only field-safe net for #148 — NO Android APK, players unchanged.** Makes the
server absorb a device that opens duplicate/rapid sockets, so a thrashing PAIRED device
converges to ONE stable connection and stays online. It does **NOT** fix the client opening
duplicate sockets (the APK duplicate-socket root cause — separate track); **#148 is not closed
on this alone.**
### Fixed / hardened — eviction storm (#148)
- **Per-device session-settle debounce.** When a device_id with a LIVE incumbent socket opens
another socket within a short window (`SESSION_SETTLE_WINDOW_MS`, default 2500ms), the
duplicate is **soft-refused and the incumbent kept** — so a duplicate burst converges on one
connection and the device stays online, instead of churning through evictions. This closes
the gap the reconnect-throttle's **30s post-restart warm-up** leaves open (during warm-up only
the hard ceiling applies, so a burst passed undamped and each new socket evicted the prior).
The debounce is **warm-up-independent**.
- **Liveness safeguard:** the incumbent is only kept if its socket is genuinely live; a
dead/half-open incumbent is replaced — the device is **never stranded offline**.
- **Soft refusal, never a quarantine** (paired-safe); single-session enforcement intact for a
legitimate move; unpaired/abusive flapping still caught by the existing limiters.
Operational note: a chunk of the observed churn was the warm-up window **re-opening on every
rapid patch redeploy** — the debounce closes that in code, but reducing redeploy frequency
independently reduces warm-up-window exposure.
Server/CMS only; ships no APK (versionCode still increments so a future player build is
OTA-recognized). Docker: `ghcr.io/screentinker/screentinker:1.9.2-patch2` (pre-release —
`:latest` stays at 1.9.2).
## 1.9.2-patch1
**Server/CMS-only connection-lifecycle hardening for #148 — NO Android APK, players stay on
their current builds.** This strictly HELPS and de-risks, but is **NOT a guaranteed #148 fix**:
the MAXHUB client-side reconnect failure and the disconnect synchronizer (edge conntrack /
reporting) are separate, unproven-here tracks that may still require a client update / a Bold
Sophos-edge review — **do not consider #148 fully closed on this patch alone.**
### Fixed / hardened — connection lifecycle (#148)
- **The flap-limiter no longer quarantines legitimate PAIRED devices on reconnect churn.** A
paired + authenticated device reconnecting is exempt from the 30-min quarantine escalation
(a brief soft cooldown at most), so a repeated edge/NAT flush behind one SNAT IP can no
longer be amplified into a self-inflicted fleet-wide lockout. Unpaired/abusive flapping is
still quarantined (the attacker / unprovisioned-hammering case is unchanged).
- **Marking a device offline now also closes its socket**, so DB-offline can't diverge from
socket-state into a silent half-open the client is never told about.
- **Faster half-open detection:** ping interval 30s → 15s (the pong TIMEOUT is kept at 30s so
decode-loaded TV WebKits aren't falsely dropped) → dead-peer detection 60s → 45s on BOTH the
server AND the client (the client inherits these via the handshake — **no APK needed**).
- **TCP SO_KEEPALIVE** on every connection so a half-open TCP can't persist indefinitely at the
OS layer.
Server/CMS version only; ships no APK (versionCode still increments so a future player build is
OTA-recognized). Docker: `ghcr.io/screentinker/screentinker:1.9.2-patch1` (pre-release —
`:latest` stays at 1.9.2).
## 1.9.2
**⚠ Major internal hardening release (the "#146" rewrite) — large blast radius.** 1.9.2
rewrites the connection / maintenance / OTA hot paths to kill an event-loop death spiral,
plus adds usage-metering (billing) and web-player fixes. If you bisect a regression to the
1.9.x line, 1.9.2 is the big one. Core invariant introduced: **no synchronous op may block
the event loop for more than ~50ms**, ever. Every new subsystem has an env kill-switch.
### Fixed — maintenance / prune (the death-spiral root cause)
- **Non-blocking, chunked, per-device `device_status_log` prune.** The old whole-table
`ROW_NUMBER` sort froze boot for 4048s at ~1M rows → healthcheck fail → restart loop that
wiped in-memory throttle state → the spiral. Prune is now per-device, indexed, batched with
`setImmediate` yields (`lib/chunked-prune.js`), async, re-entrant, and band-gated on the
interval run (the startup prune is intentionally un-gated so a bloated table self-heals on
first boot without freezing it). All table-growth sweeps (status-log, play-logs,
provisioning, telemetry, lag) route through the chunked helper. New index
`idx_devices_provisioning`. **Measured worst-case event-loop gap under the storm harness:
<300ms across 300k rows (was 4048s).**
### Fixed — reconnect / flap
- **Per-device flap-rate limiter** (`lib/flap-limiter.js`): a device reconnecting faster than
`CONNECT_RATE_MAX` (20) per `CONNECT_RATE_WINDOW_MS` (5min) is refused at the register gate,
keyed via a **SNAT-safe identity chain** (device_id fingerprint token one bounded
global anon bucket) **never by IP** (the whole fleet egresses one IP). After repeated
trips a hard flapper is **quarantined IN-MEMORY for 30min and auto-clears** it is NOT a
durable DB block.
- **Operator block kill-switch:** `POST /api/devices/:id/{block,unblock}` + a dashboard
button; the block check resolves the effective device_id via the identity chain so a
device_id-less reconnect of a blocked device is still caught. Takes effect on next register,
no restart.
- Also folded in: false-offline fixes (live-socket liveness beats a lagged heartbeat clock;
evicted-socket re-arm race) and per-connection fail-fast so one device's handler throw can
never exit the process.
### Fixed — OTA (SNAT-safe)
- `/api/update/check` early-returns before any filesystem call when there's no offer; APK
metadata is cached. `/download/apk` gains a **band-aware** global concurrency + rate guard
that sheds with **503 Retry-After only under elevated/critical loop-lag** under normal
band, downloads serve freely (a coordinated fleet rollout is never staggered when healthy).
All limiting is global/aggregate **no per-IP limiting** (SNAT).
### Added — telemetry / logging / observability
- Batched `event_loop_lag` inserts (buffered, flushed every 10s) and coalesced high-frequency
logging (one summarized line per key per 30s; band *changes* stay immediate).
- **Throughput counters** (running total + last-completed-window) in the `/api/status` debug
block so a flapper/flood shows on the server itself (`flap.refusedLastWindow` climbing while
`band=normal` = the limiter absorbing it cheaply). The debug block is now **admin-toggleable**
(Admin tab, persisted, no restart; default follows `STATUS_DEBUG_ENABLED`).
- **`devices_connected`** on `/api/status` (always-on): the live WS-socket count from the
heartbeat connection map (NOT the lagging `devices.status='online'` column).
### Added — billing (usage metering)
- **Billable Screens** metering per the ByteTinkerBold agreement the contractual
system-of-record. A durable daily rollup (`device_usage_daily`) is accumulated incrementally
off the heartbeat tick from live presence (retention-independent), pruned chunked. Exposed on
a **dedicated, admin-gated `GET /api/billing/usage`** route (NOT on `/api/status`; billing is
revenue data). Readable via an **owner-minted, revocable `billing:read` scoped token**
(`scripts/mint-billing-token.js`) that authorizes billing-read and nothing else, OR a
platform-admin session. See [`docs/billing.md`](docs/billing.md).
### Fixed — web player
- **"Unchanged" refresh no longer drops the video.** On a reconnect the server re-emits
`device:paired` while content is already playing; the player showed the idle "Waiting for
content…" overlay unconditionally (covering live video; audio kept playing underneath) and
the following "Playlist unchanged" left it up. Idle now shows only when genuinely idle, and
an unchanged refresh is a strict no-op that leaves playback exactly as-is.
- Hardened `PlayerMediaHealth` call sites to guard by **method** (not object) so a stale-cached
player module can't throw `shouldShowIdle is not a function` and abort a socket handler.
### Added — translations
- Italian (`it`) locale updated (#145).
## 1.9.2-beta1 — unreleased
### Fixed — server resilience (#142)
- **A single flapping device can no longer saturate the event loop.** A new
load-aware, per-device reconnect throttle (`lib/reconnect-throttle.js`) gates
genuine reconnects *before* the heavy register work (DB writes + playlist build).
The verdict is per-device; global event-loop lag only multiplies an
already-flagged device's backoff and never throttles a healthy one. Hard ceiling
+ cold-start warm-up so a full-fleet reconnect after a deploy is never throttled.
- **`device_status_log` growth is bounded.** Added
`idx_device_status_log_device_ts`, a global retention sweep (`pruneStatusLog`,
`STATUS_LOG_RETENTION_DAYS` default 3) covering removed/idle devices and the
`offline_timeout` path, and de-duplicated the table's `CREATE TABLE`.
- **`content-ack` spam de-duplicated.** Repeated identical
`(device_id, content_id, status)` reports are suppressed within
`CONTENT_ACK_DEDUP_MS` (default 10s).
- **Provisioning cleanup window corrected.** Unclaimed provisioning devices are now
swept after 24h (the code used `365 * 86400` a year contradicting its own
comment).
### Added — observability (#142)
- **Event-loop lag telemetry** via `perf_hooks.monitorEventLoopDelay()`. Sampled to
a bounded `event_loop_lag` table (indexed + pruned, `LAG_TELEMETRY_RETENTION_DAYS`)
and surfaced on `/api/status` as `loop_lag` (mean/p50/p99/max + band).
### Maintenance
- Operators whose `device_status_log` is already bloated from a pre-1.9.2 deployment
should reclaim disk with a **one-time manual `VACUUM`** in a maintenance window;
retention now bounds further growth. Auto-VACUUM is intentionally not enabled.
See [`docs/maintenance-device-status-log.md`](docs/maintenance-device-status-log.md).
## 1.9.1-beta3 — unreleased
### Fixed — Tizen player
- **#118 Sticky "Not authenticated" banner.** On TV sleep/wake the socket reconnects and
a heartbeat could fire on the fresh, not-yet-registered socket; the server rejected it
with `device:auth-error`, which the player showed as a *sticky* toast over still-playing
content (and, worse, dropped its saved credentials and re-paired). Heartbeats are now
gated on a per-connection `authenticated` flag (set only between `device:registered` and
`disconnect`/`auth-error`), the heartbeat timer is stopped on `connect`/`disconnect`/
`auth-error`, the stale banner is cleared on `device:registered`, and the `auth-error`
toast is non-sticky so any transient case self-clears.
- **#119 `app_version` stuck at `1.0.0`.** The hardcoded constant made every Tizen device
report `1.0.0` regardless of the installed `.wgt`. The version now resolves at runtime
from `config.xml` via the Tizen application API, with a fallback constant that
`build-wgt.sh` stamps from `config.xml`'s `version=""`.
### Added — Tizen player
- **Video walls (`wall:sync`).** The Tizen player now supports wall membership: when the
payload carries `wall_config`, a new `WallController` positions the stage (vw/vh) as this
screen's slice of the wall and drives the single-zone player as leader or follower. The
leader broadcasts `wall:sync` at 4Hz; followers align their index and keep their video
locked to the leader's clock with a latency-compensated drift controller (hard-seek past
0.3s, gentle ±3% playbackRate nudge past 0.05s), and request an immediate position on
(re)connect via `wall:sync-request`. Mirrors the web player (the Android player has no
wall support). Per-tile `rotation` is not applied yet (web-player parity). Wall emits are
gated on auth + connection so a pre-register tick can't trip `device:auth-error`.
- **Multi-zone layouts (Android parity).** The Tizen player now renders assigned layouts,
not just fullscreen single-zone. A new `ZoneRenderer` (ports the Android `ZoneManager`)
positions zones by percent geometry with `z_index`/`fit_mode`/background, groups
assignments by `zone_id` (unassigned content goes to the first zone), and rotates each
zone independently with the same per-item schedule gating (#74/#75). `app.js` selects the
renderer from `payload.layout`; single-zone playback is unchanged. (Video walls
`wall:sync` are still Android-only.)
- **#121 Remote commands.** Added a `device:command` handler (`refresh`, `launch`,
`screen_on`, `screen_off`, plus honest no-op toasts for `update`/`reboot`/`shutdown`,
which need B2B/MDM privileges a sideloaded app lacks). Removed the dead `device:reload`
listener (the server never emitted it) in favour of `device:command` `refresh`.
- **#120 Dashboard preview.** Added `device:screenshot-request` / `device:remote-start` /
`device:remote-stop`. Images capture for real; `<video>`/YouTube fall back to a status
card because the TV's hardware video plane and cross-origin iframes can't be read into a
`<canvas>`. See `tizen/README.md` for the support matrix.
- **#122 Updates / boot.** Documented the supported paths `.wgt` re-sideload or URL
Launcher/MDM refresh for updates, and display-level kiosk/URL-Launcher settings for
auto-launch on boot (there is no in-app OTA or `config.xml` autostart for a sideloaded
consumer TV web app).
## 1.9.0 — 2026-06-11
### Added
- **Per-playlist-item schedules.** Each playlist item can carry one or more schedule
blocks active days, a start/end time-of-day, and optional start/end dates. An item
plays when the screen's local "now" matches at least one block; an item with no
blocks always plays. Edit per item via the clock icon in the playlist editor (a badge
summarises the schedule on each row).
- **#74 dayparting:** time-of-day + day-of-week windows, including overnight windows
that cross midnight (a Fri 22:0002:00 block is active Sat 01:00).
- **#75 auto-expire:** inclusive start/end dates; an item past its end date stops
showing automatically even on offline screens, because evaluation is on-device.
- All three players (web, Android, Tizen) evaluate schedules client-side against their
own clock, so dayparting and expiry work offline. They share one evaluator contract,
`shared/schedule-vectors.json` 39 conformance vectors covering DST (US + AU),
overnight-wrap day anchoring, timezone correctness, and date boundaries. CI runs the
vectors against the JS evaluator (node) and the Kotlin port (Gradle/JUnit); the Tizen
copy is byte-identical to the JS source and checked under node.
- Device detail now shows the screen's reported timezone and clock, with a **clock-skew
warning** when the device clock differs from the server by more than 2 minutes (a bad
device clock makes schedules fire at the wrong local time).
### Changed — device-level schedule timezone (behaviour change)
- Device/group **schedule overrides** (the existing calendar feature) are now evaluated
in each device's effective timezone instead of the server's local time. Previously the
`schedules.timezone` field was never applied and "07:00" meant the *server's* 07:00.
Now "07:00" means the *screen's* 07:00 which is what was intended.
- **Who is affected:** self-hosters whose server timezone differs from their screens'
timezone their existing device schedules will shift to fire at the screens' local
time. Single-timezone deployments (server and screens in the same zone) are
unaffected. A device with no timezone set and not reporting one falls back to the
server clock (unchanged from before).
### Fixed
- **#81 release APK is now v1 + v2 + v3 signed.** With `minSdk 26`, the Android Gradle
Plugin defaulted the v1 (JAR) signature *off*, producing a v2-only APK that some
MDM-managed commercial signage (e.g. MAXHUB via the Pivot MDM) silently removes on the
next reboot so screens that power-cycle nightly lost the app and fell back to the
setup screen. Setting `enableV1Signing = true` had no effect at minSdk 24; the release
build now re-signs with `apksigner` and a low `--min-sdk-version` to emit the JAR
signature alongside v2/v3. Verified to install and run on Android 14+/API 36 as well.
### Notes
- **Scheduling fails open.** If the on-device evaluator ever errors (bad timezone id,
malformed block), the item **plays** rather than being hidden. A blank screen is worse
than an over-running promo this is a guarantee, enforced in all three players.
- Windows are enforced at **item boundaries**: a long item finishes before the schedule
is re-checked, so it can overshoot its window by up to its own duration.
- **A single video *with a schedule* now re-renders at each loop boundary** so its window
can be re-evaluated; seamless native looping still applies to unscheduled single videos.
Deliberate tradeoff a brief seam each loop for a scheduled lone video, in exchange for
its daypart/expiry actually being honoured.
- **Re-publish required:** editing a schedule puts the playlist into draft; publish to
push schedules to devices. Existing published playlists keep playing unchanged until
re-published.
- Players that predate this release ignore the new fields and keep playing everything
(graceful degradation) update players to honour schedules.