mirror of
https://github.com/screentinker/screentinker.git
synced 2026-08-13 22:03:13 -06:00
265 lines
17 KiB
Markdown
265 lines
17 KiB
Markdown
# Changelog
|
||
|
||
## 1.9.2-patch2
|
||
|
||
**Server/CMS-only field-safe net for #148 — NO Android APK, players unchanged.** Makes the
|
||
server absorb a device that opens duplicate/rapid sockets, so a thrashing PAIRED device
|
||
converges to ONE stable connection and stays online. It does **NOT** fix the client opening
|
||
duplicate sockets (the APK duplicate-socket root cause — separate track); **#148 is not closed
|
||
on this alone.**
|
||
|
||
### Fixed / hardened — eviction storm (#148)
|
||
- **Per-device session-settle debounce.** When a device_id with a LIVE incumbent socket opens
|
||
another socket within a short window (`SESSION_SETTLE_WINDOW_MS`, default 2500ms), the
|
||
duplicate is **soft-refused and the incumbent kept** — so a duplicate burst converges on one
|
||
connection and the device stays online, instead of churning through evictions. This closes
|
||
the gap the reconnect-throttle's **30s post-restart warm-up** leaves open (during warm-up only
|
||
the hard ceiling applies, so a burst passed undamped and each new socket evicted the prior).
|
||
The debounce is **warm-up-independent**.
|
||
- **Liveness safeguard:** the incumbent is only kept if its socket is genuinely live; a
|
||
dead/half-open incumbent is replaced — the device is **never stranded offline**.
|
||
- **Soft refusal, never a quarantine** (paired-safe); single-session enforcement intact for a
|
||
legitimate move; unpaired/abusive flapping still caught by the existing limiters.
|
||
|
||
Operational note: a chunk of the observed churn was the warm-up window **re-opening on every
|
||
rapid patch redeploy** — the debounce closes that in code, but reducing redeploy frequency
|
||
independently reduces warm-up-window exposure.
|
||
|
||
Server/CMS only; ships no APK (versionCode still increments so a future player build is
|
||
OTA-recognized). Docker: `ghcr.io/screentinker/screentinker:1.9.2-patch2` (pre-release —
|
||
`:latest` stays at 1.9.2).
|
||
|
||
## 1.9.2-patch1
|
||
|
||
**Server/CMS-only connection-lifecycle hardening for #148 — NO Android APK, players stay on
|
||
their current builds.** This strictly HELPS and de-risks, but is **NOT a guaranteed #148 fix**:
|
||
the MAXHUB client-side reconnect failure and the disconnect synchronizer (edge conntrack /
|
||
reporting) are separate, unproven-here tracks that may still require a client update / a Bold
|
||
Sophos-edge review — **do not consider #148 fully closed on this patch alone.**
|
||
|
||
### Fixed / hardened — connection lifecycle (#148)
|
||
- **The flap-limiter no longer quarantines legitimate PAIRED devices on reconnect churn.** A
|
||
paired + authenticated device reconnecting is exempt from the 30-min quarantine escalation
|
||
(a brief soft cooldown at most), so a repeated edge/NAT flush behind one SNAT IP can no
|
||
longer be amplified into a self-inflicted fleet-wide lockout. Unpaired/abusive flapping is
|
||
still quarantined (the attacker / unprovisioned-hammering case is unchanged).
|
||
- **Marking a device offline now also closes its socket**, so DB-offline can't diverge from
|
||
socket-state into a silent half-open the client is never told about.
|
||
- **Faster half-open detection:** ping interval 30s → 15s (the pong TIMEOUT is kept at 30s so
|
||
decode-loaded TV WebKits aren't falsely dropped) → dead-peer detection 60s → 45s on BOTH the
|
||
server AND the client (the client inherits these via the handshake — **no APK needed**).
|
||
- **TCP SO_KEEPALIVE** on every connection so a half-open TCP can't persist indefinitely at the
|
||
OS layer.
|
||
|
||
Server/CMS version only; ships no APK (versionCode still increments so a future player build is
|
||
OTA-recognized). Docker: `ghcr.io/screentinker/screentinker:1.9.2-patch1` (pre-release —
|
||
`:latest` stays at 1.9.2).
|
||
|
||
## 1.9.2
|
||
|
||
**⚠ Major internal hardening release (the "#146" rewrite) — large blast radius.** 1.9.2
|
||
rewrites the connection / maintenance / OTA hot paths to kill an event-loop death spiral,
|
||
plus adds usage-metering (billing) and web-player fixes. If you bisect a regression to the
|
||
1.9.x line, 1.9.2 is the big one. Core invariant introduced: **no synchronous op may block
|
||
the event loop for more than ~50ms**, ever. Every new subsystem has an env kill-switch.
|
||
|
||
### Fixed — maintenance / prune (the death-spiral root cause)
|
||
- **Non-blocking, chunked, per-device `device_status_log` prune.** The old whole-table
|
||
`ROW_NUMBER` sort froze boot for 40–48s at ~1M rows → healthcheck fail → restart loop that
|
||
wiped in-memory throttle state → the spiral. Prune is now per-device, indexed, batched with
|
||
`setImmediate` yields (`lib/chunked-prune.js`), async, re-entrant, and band-gated on the
|
||
interval run (the startup prune is intentionally un-gated so a bloated table self-heals on
|
||
first boot without freezing it). All table-growth sweeps (status-log, play-logs,
|
||
provisioning, telemetry, lag) route through the chunked helper. New index
|
||
`idx_devices_provisioning`. **Measured worst-case event-loop gap under the storm harness:
|
||
<300ms across 300k rows (was 40–48s).**
|
||
|
||
### Fixed — reconnect / flap
|
||
- **Per-device flap-rate limiter** (`lib/flap-limiter.js`): a device reconnecting faster than
|
||
`CONNECT_RATE_MAX` (20) per `CONNECT_RATE_WINDOW_MS` (5min) is refused at the register gate,
|
||
keyed via a **SNAT-safe identity chain** (device_id → fingerprint → token → one bounded
|
||
global anon bucket) — **never by IP** (the whole fleet egresses one IP). After repeated
|
||
trips a hard flapper is **quarantined IN-MEMORY for 30min and auto-clears** — it is NOT a
|
||
durable DB block.
|
||
- **Operator block kill-switch:** `POST /api/devices/:id/{block,unblock}` + a dashboard
|
||
button; the block check resolves the effective device_id via the identity chain so a
|
||
device_id-less reconnect of a blocked device is still caught. Takes effect on next register,
|
||
no restart.
|
||
- Also folded in: false-offline fixes (live-socket liveness beats a lagged heartbeat clock;
|
||
evicted-socket re-arm race) and per-connection fail-fast so one device's handler throw can
|
||
never exit the process.
|
||
|
||
### Fixed — OTA (SNAT-safe)
|
||
- `/api/update/check` early-returns before any filesystem call when there's no offer; APK
|
||
metadata is cached. `/download/apk` gains a **band-aware** global concurrency + rate guard
|
||
that sheds with **503 Retry-After only under elevated/critical loop-lag** — under normal
|
||
band, downloads serve freely (a coordinated fleet rollout is never staggered when healthy).
|
||
All limiting is global/aggregate — **no per-IP limiting** (SNAT).
|
||
|
||
### Added — telemetry / logging / observability
|
||
- Batched `event_loop_lag` inserts (buffered, flushed every 10s) and coalesced high-frequency
|
||
logging (one summarized line per key per 30s; band *changes* stay immediate).
|
||
- **Throughput counters** (running total + last-completed-window) in the `/api/status` debug
|
||
block so a flapper/flood shows on the server itself (`flap.refusedLastWindow` climbing while
|
||
`band=normal` = the limiter absorbing it cheaply). The debug block is now **admin-toggleable**
|
||
(Admin tab, persisted, no restart; default follows `STATUS_DEBUG_ENABLED`).
|
||
- **`devices_connected`** on `/api/status` (always-on): the live WS-socket count from the
|
||
heartbeat connection map (NOT the lagging `devices.status='online'` column).
|
||
|
||
### Added — billing (usage metering)
|
||
- **Billable Screens** metering per the ByteTinker–Bold agreement — the contractual
|
||
system-of-record. A durable daily rollup (`device_usage_daily`) is accumulated incrementally
|
||
off the heartbeat tick from live presence (retention-independent), pruned chunked. Exposed on
|
||
a **dedicated, admin-gated `GET /api/billing/usage`** route (NOT on `/api/status`; billing is
|
||
revenue data). Readable via an **owner-minted, revocable `billing:read` scoped token**
|
||
(`scripts/mint-billing-token.js`) that authorizes billing-read and nothing else, OR a
|
||
platform-admin session. See [`docs/billing.md`](docs/billing.md).
|
||
|
||
### Fixed — web player
|
||
- **"Unchanged" refresh no longer drops the video.** On a reconnect the server re-emits
|
||
`device:paired` while content is already playing; the player showed the idle "Waiting for
|
||
content…" overlay unconditionally (covering live video; audio kept playing underneath) and
|
||
the following "Playlist unchanged" left it up. Idle now shows only when genuinely idle, and
|
||
an unchanged refresh is a strict no-op that leaves playback exactly as-is.
|
||
- Hardened `PlayerMediaHealth` call sites to guard by **method** (not object) so a stale-cached
|
||
player module can't throw `shouldShowIdle is not a function` and abort a socket handler.
|
||
|
||
### Added — translations
|
||
- Italian (`it`) locale updated (#145).
|
||
|
||
|
||
## 1.9.2-beta1 — unreleased
|
||
|
||
### Fixed — server resilience (#142)
|
||
- **A single flapping device can no longer saturate the event loop.** A new
|
||
load-aware, per-device reconnect throttle (`lib/reconnect-throttle.js`) gates
|
||
genuine reconnects *before* the heavy register work (DB writes + playlist build).
|
||
The verdict is per-device; global event-loop lag only multiplies an
|
||
already-flagged device's backoff and never throttles a healthy one. Hard ceiling
|
||
+ cold-start warm-up so a full-fleet reconnect after a deploy is never throttled.
|
||
- **`device_status_log` growth is bounded.** Added
|
||
`idx_device_status_log_device_ts`, a global retention sweep (`pruneStatusLog`,
|
||
`STATUS_LOG_RETENTION_DAYS` default 3) covering removed/idle devices and the
|
||
`offline_timeout` path, and de-duplicated the table's `CREATE TABLE`.
|
||
- **`content-ack` spam de-duplicated.** Repeated identical
|
||
`(device_id, content_id, status)` reports are suppressed within
|
||
`CONTENT_ACK_DEDUP_MS` (default 10s).
|
||
- **Provisioning cleanup window corrected.** Unclaimed provisioning devices are now
|
||
swept after 24h (the code used `365 * 86400` — a year — contradicting its own
|
||
comment).
|
||
|
||
### Added — observability (#142)
|
||
- **Event-loop lag telemetry** via `perf_hooks.monitorEventLoopDelay()`. Sampled to
|
||
a bounded `event_loop_lag` table (indexed + pruned, `LAG_TELEMETRY_RETENTION_DAYS`)
|
||
and surfaced on `/api/status` as `loop_lag` (mean/p50/p99/max + band).
|
||
|
||
### Maintenance
|
||
- Operators whose `device_status_log` is already bloated from a pre-1.9.2 deployment
|
||
should reclaim disk with a **one-time manual `VACUUM`** in a maintenance window;
|
||
retention now bounds further growth. Auto-VACUUM is intentionally not enabled.
|
||
See [`docs/maintenance-device-status-log.md`](docs/maintenance-device-status-log.md).
|
||
|
||
## 1.9.1-beta3 — unreleased
|
||
|
||
### Fixed — Tizen player
|
||
- **#118 Sticky "Not authenticated" banner.** On TV sleep/wake the socket reconnects and
|
||
a heartbeat could fire on the fresh, not-yet-registered socket; the server rejected it
|
||
with `device:auth-error`, which the player showed as a *sticky* toast over still-playing
|
||
content (and, worse, dropped its saved credentials and re-paired). Heartbeats are now
|
||
gated on a per-connection `authenticated` flag (set only between `device:registered` and
|
||
`disconnect`/`auth-error`), the heartbeat timer is stopped on `connect`/`disconnect`/
|
||
`auth-error`, the stale banner is cleared on `device:registered`, and the `auth-error`
|
||
toast is non-sticky so any transient case self-clears.
|
||
- **#119 `app_version` stuck at `1.0.0`.** The hardcoded constant made every Tizen device
|
||
report `1.0.0` regardless of the installed `.wgt`. The version now resolves at runtime
|
||
from `config.xml` via the Tizen application API, with a fallback constant that
|
||
`build-wgt.sh` stamps from `config.xml`'s `version=""`.
|
||
|
||
### Added — Tizen player
|
||
- **Video walls (`wall:sync`).** The Tizen player now supports wall membership: when the
|
||
payload carries `wall_config`, a new `WallController` positions the stage (vw/vh) as this
|
||
screen's slice of the wall and drives the single-zone player as leader or follower. The
|
||
leader broadcasts `wall:sync` at 4Hz; followers align their index and keep their video
|
||
locked to the leader's clock with a latency-compensated drift controller (hard-seek past
|
||
0.3s, gentle ±3% playbackRate nudge past 0.05s), and request an immediate position on
|
||
(re)connect via `wall:sync-request`. Mirrors the web player (the Android player has no
|
||
wall support). Per-tile `rotation` is not applied yet (web-player parity). Wall emits are
|
||
gated on auth + connection so a pre-register tick can't trip `device:auth-error`.
|
||
- **Multi-zone layouts (Android parity).** The Tizen player now renders assigned layouts,
|
||
not just fullscreen single-zone. A new `ZoneRenderer` (ports the Android `ZoneManager`)
|
||
positions zones by percent geometry with `z_index`/`fit_mode`/background, groups
|
||
assignments by `zone_id` (unassigned content goes to the first zone), and rotates each
|
||
zone independently with the same per-item schedule gating (#74/#75). `app.js` selects the
|
||
renderer from `payload.layout`; single-zone playback is unchanged. (Video walls
|
||
`wall:sync` are still Android-only.)
|
||
- **#121 Remote commands.** Added a `device:command` handler (`refresh`, `launch`,
|
||
`screen_on`, `screen_off`, plus honest no-op toasts for `update`/`reboot`/`shutdown`,
|
||
which need B2B/MDM privileges a sideloaded app lacks). Removed the dead `device:reload`
|
||
listener (the server never emitted it) in favour of `device:command` `refresh`.
|
||
- **#120 Dashboard preview.** Added `device:screenshot-request` / `device:remote-start` /
|
||
`device:remote-stop`. Images capture for real; `<video>`/YouTube fall back to a status
|
||
card because the TV's hardware video plane and cross-origin iframes can't be read into a
|
||
`<canvas>`. See `tizen/README.md` for the support matrix.
|
||
- **#122 Updates / boot.** Documented the supported paths — `.wgt` re-sideload or URL
|
||
Launcher/MDM refresh for updates, and display-level kiosk/URL-Launcher settings for
|
||
auto-launch on boot (there is no in-app OTA or `config.xml` autostart for a sideloaded
|
||
consumer TV web app).
|
||
|
||
## 1.9.0 — 2026-06-11
|
||
|
||
### Added
|
||
- **Per-playlist-item schedules.** Each playlist item can carry one or more schedule
|
||
blocks — active days, a start/end time-of-day, and optional start/end dates. An item
|
||
plays when the screen's local "now" matches at least one block; an item with no
|
||
blocks always plays. Edit per item via the clock icon in the playlist editor (a badge
|
||
summarises the schedule on each row).
|
||
- **#74 dayparting:** time-of-day + day-of-week windows, including overnight windows
|
||
that cross midnight (a Fri 22:00–02:00 block is active Sat 01:00).
|
||
- **#75 auto-expire:** inclusive start/end dates; an item past its end date stops
|
||
showing automatically — even on offline screens, because evaluation is on-device.
|
||
- All three players (web, Android, Tizen) evaluate schedules client-side against their
|
||
own clock, so dayparting and expiry work offline. They share one evaluator contract,
|
||
`shared/schedule-vectors.json` — 39 conformance vectors covering DST (US + AU),
|
||
overnight-wrap day anchoring, timezone correctness, and date boundaries. CI runs the
|
||
vectors against the JS evaluator (node) and the Kotlin port (Gradle/JUnit); the Tizen
|
||
copy is byte-identical to the JS source and checked under node.
|
||
- Device detail now shows the screen's reported timezone and clock, with a **clock-skew
|
||
warning** when the device clock differs from the server by more than 2 minutes (a bad
|
||
device clock makes schedules fire at the wrong local time).
|
||
|
||
### Changed — device-level schedule timezone (behaviour change)
|
||
- Device/group **schedule overrides** (the existing calendar feature) are now evaluated
|
||
in each device's effective timezone instead of the server's local time. Previously the
|
||
`schedules.timezone` field was never applied and "07:00" meant the *server's* 07:00.
|
||
Now "07:00" means the *screen's* 07:00 — which is what was intended.
|
||
- **Who is affected:** self-hosters whose server timezone differs from their screens'
|
||
timezone — their existing device schedules will shift to fire at the screens' local
|
||
time. Single-timezone deployments (server and screens in the same zone) are
|
||
unaffected. A device with no timezone set and not reporting one falls back to the
|
||
server clock (unchanged from before).
|
||
|
||
### Fixed
|
||
- **#81 — release APK is now v1 + v2 + v3 signed.** With `minSdk 26`, the Android Gradle
|
||
Plugin defaulted the v1 (JAR) signature *off*, producing a v2-only APK that some
|
||
MDM-managed commercial signage (e.g. MAXHUB via the Pivot MDM) silently removes on the
|
||
next reboot — so screens that power-cycle nightly lost the app and fell back to the
|
||
setup screen. Setting `enableV1Signing = true` had no effect at minSdk ≥ 24; the release
|
||
build now re-signs with `apksigner` and a low `--min-sdk-version` to emit the JAR
|
||
signature alongside v2/v3. Verified to install and run on Android 14+/API 36 as well.
|
||
|
||
### Notes
|
||
- **Scheduling fails open.** If the on-device evaluator ever errors (bad timezone id,
|
||
malformed block), the item **plays** rather than being hidden. A blank screen is worse
|
||
than an over-running promo — this is a guarantee, enforced in all three players.
|
||
- Windows are enforced at **item boundaries**: a long item finishes before the schedule
|
||
is re-checked, so it can overshoot its window by up to its own duration.
|
||
- **A single video *with a schedule* now re-renders at each loop boundary** so its window
|
||
can be re-evaluated; seamless native looping still applies to unscheduled single videos.
|
||
Deliberate tradeoff — a brief seam each loop for a scheduled lone video, in exchange for
|
||
its daypart/expiry actually being honoured.
|
||
- **Re-publish required:** editing a schedule puts the playlist into draft; publish to
|
||
push schedules to devices. Existing published playlists keep playing unchanged until
|
||
re-published.
|
||
- Players that predate this release ignore the new fields and keep playing everything
|
||
(graceful degradation) — update players to honour schedules.
|