Commit graph

6 commits

Author SHA1 Message Date
ScreenTinker b25bfaea57 Stop the player rendering black when the server hosts the display
A BrightSign hosting ScreenTinker shows a local page from file:///ssd:/node-server.html
that layers the player in an iframe — an iframe rather than a navigation, because
navigating replaces the document and kills the poller that notices the server dying.

helmet sets X-Frame-Options: SAMEORIGIN, and file:// is not the same origin as
http://127.0.0.1:8181, so the frame rendered BLACK. Every asset inside it returned 200
— the player page and all six of its scripts — and nothing appeared in any log. Only
the response headers said why, which is a miserable thing to debug on a device with no
console.

⚠️ AND IT IS NOT ONLY /player. Chrome evaluates SAMEORIGIN against the TOP-LEVEL
document rather than the immediate parent, so with a file:// page at the top, every
iframe the player itself uses — widget renders, kiosk views — is blocked by the same
rule one level deeper. Scoping this to /player would have cleared the black screen and
left every widget in the playlist black instead: the same bug, found later, on a
customer's wall. The test covers that case explicitly.

Scoped by CONTEXT, not by path: only when the process was started as a player host
(bs-server-boot.js sets ST_PLAYER_HOST; nothing else does) AND the request arrived on
loopback, i.e. from the box's own browser. An ordinary server keeps SAMEORIGIN, and so
does any request off the network — which is where clickjacking would have to come from,
since a remote page cannot reach another machine's 127.0.0.1. Where a CSP is set, only
its frame-ancestors directive is rewritten; the rest of the policy survives.

Verified on XT245 URD3C6000823: the player renders and shows its pairing code, while a
request to the same URL from the LAN still returns X-Frame-Options: SAMEORIGIN.
Full suite 1766 pass / 0 fail.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014kfhrUPit5MCqxeTQyqr56
2026-08-18 21:36:56 -05:00
ScreenTinker 4a4b9e343a brightsign: stage our own ffprobe/ffmpeg into /tmp so media tools work
The server logged "[MEDIA] ffmpeg, ffprobe not found on PATH — video thumbnails
and durations are DISABLED" on every boot of a player-hosted server. It now says
"found — video thumbnails enabled", because the binaries are shipped gzipped,
unpacked into /tmp at startup and put on PATH before server.js is required (its
probe looks them up by name with execFile, so PATH is the whole mechanism).

/tmp is not laziness, it is the only option, and the alternatives were measured on
an XT245 rather than assumed:

  /storage/ssd    bsexfat  rw,nosuid,nodev,noexec,...     <- exec => EACCES
  /storage/flash  ext4     rw,nosuid,nodev,noexec,...     <- same
  /storage/tmp    tmpfs    rw,nosuid,nodev,noexec,...     <- same
  /tmp            tmpfs    rw,relatime                    <- the one that permits exec

We run as uid=994(nodejs), so `mount -o remount,exec` answers "permission denied
(are you root?)" on both volumes, and there is no setuid path to it: BrightSign
points mount at busybox.nosuid, and their busybox.suid carries login/passwd/vlock
and no mount applet. A symlink does not help either — noexec is a property of the
filesystem holding the inode, not of the path used to reach it, so a link in /tmp
pointing at flash still fails EACCES. Copying is what moves the inode onto a
filesystem that permits execution.

The binaries are ours and deliberately link nothing of BrightSign's. The OS does
ship the whole ffmpeg 5.1 stack (libavformat/libavcodec/... backing GStreamer) and
a stock Debian ffprobe against those libs starts, prints its banner, and then
SIGSEGVs the moment it opens a file — their Yocto build is patched for hardware
decode. So these are cross-built FFmpeg 7.1.1, --disable-gpl (LGPL 2.1+), fully
static, --enable-small. ffprobe carries no decoders at all (durations and geometry
come from the container) which is why it is 1.8MB against ffmpeg's 4.9MB; ~6.7MB of
tmpfs on a box with 2.8GB free, ~3.2MB gzipped at rest.

Verified on XT245 URD3C6000823: unpack 19ms, ffprobe -version 9ms,
format=duration 3.533333 against a 3.533333s clip, stream 320x240 h264.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014kfhrUPit5MCqxeTQyqr56
2026-08-18 20:12:26 -05:00
ScreenTinker a87bd7d875 brightsign: say what the 8182 status listener is for on screen
"status listener on 127.0.0.1:8182" told an operator a port was open and nothing
about why, next to a server that advertises 8181 - it reads like a stray listener.
It now says it feeds the diagnostics screen while the app is downloading, starting
or down, and that nothing off-device can reach it.

The port is interpolated through currentPort() rather than captured: server.env is
read after this module is evaluated, so a captured value would print the 3001
default instead of the port actually serving. Verified on the device - the line
renders ":8181".

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014kfhrUPit5MCqxeTQyqr56
2026-08-18 17:42:34 -05:00
screentinker 0b063cd415
Make the on-device server opt-in, and stop the status port answering the LAN (#291)
Some checks are pending
CI / Unit tests (node --test) (push) Waiting to run
CI / OpenAPI spec lint (push) Waiting to run
CI / Android unit tests (Kotlin schedule evaluator vectors) (push) Waiting to run
CI / Licence gate + SBOM (production deps) (push) Waiting to run
CI / Boot smoke + version check (push) Waiting to run
A fleet gets one package, and exactly one box per site should host the server.
Defaulting to on would mean every player that ever received this package
started listening on 8181, and the mistake would stay invisible until two of
them fought over the same displays.

st-config.json on the storage root, {"server": 1}, switches it on. Absent,
unreadable, unparseable, or anything other than an affirmative value leaves it
off - there is no reading of a broken config file that should end with a
device deciding to host a server. It sits at the root rather than in data/
because that is where an operator drops it over the DWS, and autozip never
writes it, so a re-provision cannot silently flip a site either way. The
package ships st-config.example.json, never st-config.json, for the same
reason.

With the server off, NOTHING listens: roNodeJs is never created, so there is
no 8181 and no 8182. The page is told through its URL rather than discovering
it, because "nothing is answering" would otherwise render as a fault and send
someone looking for a server that was never meant to exist. It now has a
fourth state that says so and offers the one line of JSON that changes it.

Separately: the status listener was bound to every interface, so anything on
the customer's LAN could read the install log, disk usage, the device's own
address and a tail of the server's console - that last one carries whatever
the server printed most recently. Its only consumer is a page on the same
device. Now 127.0.0.1 only, confirmed against /proc/net/tcp rather than by
probing, after a first attempt at verifying it fell back to loopback and
reported that as the LAN result.


Claude-Session: https://claude.ai/code/session_014kfhrUPit5MCqxeTQyqr56

Co-authored-by: Dan Walters <dan.walters@bytetinker.net>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 16:32:37 -05:00
screentinker 326da8a730
Show the player on the player, and the diagnostics when something is wrong (#290)
The box is now both server and player, so its screen has to be one or the
other at any moment. Three states, and the transitions are the point:

  installing / down / failed  diagnostics, so the fault is visible
  up, but no account yet      diagnostics plus the address to create one
  up, and an account exists   the player, full screen

A fresh install has nothing to play and nobody to play it for, so it stays on
the configuration screen until someone has signed up. Hiding that address
would leave the device unsetuppable: it has no keyboard.

⚠️ THE PLAYER IS AN IFRAME LAYER, NOT A NAVIGATION. Setting location.href
would replace the document and take the poller with it - and that poller is
the only thing able to notice the server failing later. As a layer, the
diagnostics are one style change away from being back on screen, which is
exactly what should happen when a server that has been playing for weeks
throws at 3am. A test asserts location.href is never assigned, so this cannot
be quietly simplified back.

Whether an account exists is asked by the wrapper, not the page:
/api/auth/config is public, but the page is loaded from file:// - origin
"null" - and the server sets no CORS headers on its own API, while this
process is already talking to it. The answer is three-valued. null means the
probe has not replied yet and is deliberately NOT treated as false: guessing
would flip a fresh box to an empty player and take the sign-up address off
the screen while someone was reading it.

Verified against a real server rather than by inspection - install, sign up,
watch it flip:

  BEFORE signup : needsSetup=null   -> diagnostics
  POST /api/auth/register -> HTTP 201
  AFTER  signup : needsSetup=false  -> player


Claude-Session: https://claude.ai/code/session_014kfhrUPit5MCqxeTQyqr56

Co-authored-by: Dan Walters <dan.walters@bytetinker.net>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
2026-08-18 16:23:23 -05:00
screentinker 9a1a82a100
Run the ScreenTinker server on the player it serves (#288)
* Make a BrightSign say what it is running, and what it is plugged into

A panel on a wall could not answer three questions an operator asks first:
which version am I, which page am I running, and which screen is that. All
three had answers already travelling over the socket; nothing was reading them.

VERSION. device_info.app_version was the literal '1.1.0-web' for every web
player, BrightSign included — the same string as PLAYER_VERSION, which already
travels separately as client_version. So the column carried no information at
all: a panel provisioned this morning and one running a year-old host reported
identically. app_version is now the ON-DEVICE host package, the artifact OTA
replaces and the only one here that can be stale, and PLAYER_VERSION is stamped
at serve time from VERSION rather than being a constant nobody bumped for the
whole 1.x line. No '-web' suffix: client_version is only compared for equality
today, but X.Y.Z-web is a semver PRERELEASE that sorts BELOW X.Y.Z, and this
project has been bitten by exactly that before.

The host version arrives asynchronously and can land after the page registers,
so register sends what it has and the heartbeat corrects the record — which also
catches the version changing under a live page, which is what a self-update is.

THE CARD SHOWED FOR NOBODY. The Info tab's version card sat inside the block
gated on android_version && !startsWith('Web/'). A BrightSign registers as
"Web/<ua>", so the panel that most needed a version never displayed one.

THE PAD THAT COULD NOT BE CLICKED. System View was gated on tier === 2. tier is
an Android device-owner concept, NOT NULL DEFAULT 0, written only by the APK —
so a BrightSign or Tizen panel sat at 0 forever and rendered HOME, BACK, POWER,
the D-pad and OK permanently pointer-events:none, for keys those players
genuinely handle. Greying an Android gate over a working control is the "button
that cannot work" the capability system exists to prevent, inverted. Only
Recents (KEYCODE_APP_SWITCH) and Settings are truly Android-only; those are now
the only things hidden.

THE PACKAGE POINTED AT THE WRONG SERVER. autorun.zip carried the committed
default, so a player self-updating from alpha or a self-hosted box was handed a
config pointing at screentinker.com — which surfaces as a pairing bug, miles
from the packaging code that caused it. It is now stamped with the URL it was
fetched from. The bytes therefore vary per origin, so the cache is keyed by
origin and BOTH routes derive it identically: the manifest checksum and the
served bytes must come from one buffer or every player downloads, fails
verification and retries forever.

EDID. getEdidIdentity() answers seven questions and cannot answer any others —
manufacturer, EDID version, physical size, gamma and the mode lists exist only
in the raw block, which getEdid() returns as 2048 bytes. The player ships those
on the register (identity, not a reading: it changes when someone swaps the
screen) and the SERVER parses them. That split is the point: a new field becomes
a server deploy instead of a bridge update behind a 4h CDN plus an OTA for the
host. Verified against real hardware — an XT245 with a CX101 decodes to RTK /
0x1010 / serial 1 / 2020w26 / 22x13cm, preferred 1920x1200@62, matching the
player's own DWS field for field. The odd-looking 62 is right: 168.5MHz over
2200 x 1245 is 61.5Hz, and rounding it to a nicer 60 would contradict the panel.

Also corrects two comments that had outgrown their reasoning: the BrightSign
capability baseline still explained its exclusions with "a canvas cannot read
the video plane", which native capture made obsolete, and player-parity.md
claimed the bridge is "always current" when a zone-wide Cloudflare Browser Cache
TTL had been rewriting its no-cache to max-age=14400 for months.

Every new guard is mutation-tested — the fix was reverted in the source and each
test confirmed to fail. 1676 -> 1714 tests, all green.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014kfhrUPit5MCqxeTQyqr56

* Run the ScreenTinker server on the player it serves

A BrightSign XT245 now downloads, installs and runs the server itself, with
the display showing what it is doing until it is up.

WHY IT NEEDED A NEW SHAPE

BrightSignOS cannot open a large autorun.zip. The 73MB build failed at boot
with "ZipArchive error at line 91", and the OS renamed it autorun.zip_invalid
- which is how a device that had already unpacked once came back up with no
autorun at all. The identical package cut to 32KB and five files boots fine;
paths (182 chars) and depth (8) are unremarkable, so the limit is in the
boot-time reader, not the archive. BrightSign's own notes acknowledge package
size as a problem and point at webpack; that route needs the dynamic requires
in scripts/ removed first, so instead autorun.zip carries only what starts the
process and the payload arrives over HTTP into a Node that has no such limit.
The payload can also be updated without re-provisioning the device.

WHY roNodeJs AND NOT THE WIDGET

The first version ran the server inside an roHtmlWidget with nodejs_enabled.
That is a Node context inside an Electron renderer, and it is not Node. Four
separate boot failures came out of it, each invisible to a local test because
a local test runs on real Node:

  - shebangs are not stripped, so any `#!/usr/bin/env node` file dies with
    "Failed to construct 'ContextifyScript': Invalid or unexpected token".
    Note it names no token - "#" is not one. An ESM file compiled as CJS says
    "Unexpected token 'export'" instead, which is how the two are told apart.
  - require() of an ESM-only package is unsupported, which plain Node 24
    handles. uuid 14 is ESM-only and 21 files import it.
  - setInterval is the DOM's and returns a NUMBER, so setInterval(...).unref()
    throws. Two call sites were unguarded; sixteen more were written
    defensively and had been silently not unreffing.
  - worker_threads cannot create a thread at all.

BrightSign's dev-cookbook is explicit: roNodeJs "for long running processes
like ... running a web server", roHtmlWidget "for browser-based apps". Their
cra-template examples do exactly this - server in roNodeJs, widget pointed at
localhost. It also fixes the lifecycle problem that was the original argument
against a server on this hardware: in a widget the server dies with the page,
taking an open SQLite WAL with it.

The shims for the first three are kept in the packager for now rather than
removed in the same change that moves the container, so that if something
breaks it is the move and not four simultaneous removals.

CHANGES THAT ARE NOT BRIGHTSIGN-SPECIFIC

  db/database.js, routes/status.js  fs.copyFileSync does not merely copy
    bytes: it fchmods the destination to match the source. exFAT has no
    permission bits, so the pre-migration snapshot failed with EPERM and the
    failure path called process.exit(1) - which inside a widget also killed
    the page, leaving a black screen and no diagnostic. The guard was right;
    the copy was wrong. lib/fsutil.js copies without touching mode.

  db/wal-checkpointer.js  the module already degraded correctly when its
    worker died or could not be respawned, but the FIRST spawn was not
    wrapped, so a host that cannot make threads lost the whole server rather
    than falling back to inline autocheckpoint.

  db/sqlite-compat.js  a better-sqlite3 facade over node:sqlite. With it the
    bundle contains no native code at all, which is what lets an x86_64
    laptop build a package for an aarch64 player. 1719/1719 tests pass on
    Node 24 through this shim.

The packager refuses to build if a source file is untracked (git ls-files
decides what ships, and lib/fsutil.js reached a player without shipping
alongside the code that required it), if any .node binary is present, if a
shebang survives, or if a database, upload, cert or .env is staged - the first
build of this package swept up a real 33MB database and 105MB of uploads.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_014kfhrUPit5MCqxeTQyqr56

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Co-authored-by: Dan Walters <dan.walters@bytetinker.net>
2026-08-18 15:16:09 -05:00