The README covers installing, upgrading, backing up and admin recovery -- the
happy paths. It says nothing about the parts that actually go wrong: which
deployment shape you are on and why the other shape's commands silently do
nothing, how to tell a deploy really took when a version string cannot prove it,
and the handful of traps that have each cost real time at least once.
docs/operations.md is that runbook. The load-bearing entries:
- the served APK is a bind-mounted FILE, so it must be replaced in place. mv
or cp gives the host a new inode while the container keeps serving the old
bytes, with nothing in any log to say so;
- the advertised apk_size must equal the served bytes or displays download,
reject and retry forever -- and the OTA query parameter is `version`, where
the wrong name produces a result that looks identical to a broken OTA;
- a version string does not prove new code is running, and neither does the
build hash: it covers the frontend, so a server-only change deploys with an
unchanged hash and looks exactly like a stale image;
- ownership before checkout, because a partial checkout leaves VERSION updated
while the code is the old release and no migrations ran;
- a service user with no home directory makes npm install nothing while
appearing to succeed;
- a prerelease sorts below its own release, and the Android update check
offers one to any older client on the stable channel;
- native modules are built for one Node ABI, and the mismatch presents as
hundreds of unrelated test failures rather than one clear error.
Deliberately generic: no addresses, hostnames, customer names or credentials, so
it is useful to anyone self-hosting rather than a description of one estate.
Every claim was checked against the code or the workflows rather than recalled.
11 KiB
Operations runbook
Running an instance day to day: deploying, verifying, rolling back, and the traps that have actually cost people time.
The README covers the happy paths — installing, updating, backups and admin recovery. This is the part you want at 2am, or when a deploy did not behave.
Contents
- Two deployment shapes
- Before you deploy
- Deploying: native (git + systemd)
- Deploying: Docker
- The served APK
- Verifying a deploy
- Rolling back
- Releases and version numbers
- Traps worth knowing before they bite
Two deployment shapes
An instance is either native (a git checkout on a release tag, run by systemd) or Docker (a published image, run by compose). They are not interchangeable, and the commands differ at every step.
⚠️ Know which one you are on before you type anything. The most expensive mistakes in this runbook come from applying one shape's procedure to the other —
git checkouton a Docker host changes nothing the container is running, and bumping an image tag on a native host does nothing at all. If you run more than one instance, keep a note of which is which somewhere you will read.
Before you deploy
Every time, in this order:
- Snapshot the database.
sqlite3 <db> ".backup /path/to/pre-<version>-$(date +%Y%m%d-%H%M%S).db" sqlite3 /path/to/pre-<version>-*.db "PRAGMA integrity_check;" # want: ok - Record the row counts you intend to still have afterwards:
SELECT (SELECT COUNT(*) FROM devices), (SELECT COUNT(*) FROM users), (SELECT COUNT(*) FROM content), (SELECT COUNT(*) FROM playlists); - Check whether dependencies changed. If
server/package.jsondiffers by more than the version field between the running release and the target, you need an install step. If it differs only in"version", skip it — that is the cheapest and safest kind of deploy.git diff <current-tag> <target-tag> -- server/package.json - Check whether migrations will run. They apply automatically at boot. Additive columns and new
tables are safe, and a code-only rollback simply leaves them unused.
git diff <current-tag> <target-tag> -- server/db/database.js | grep -E '^\+.*(ALTER|CREATE) TABLE|CREATE INDEX' - Back up the compose file / the served APK if you are about to change either.
Deploying: native (git + systemd)
scripts/upgrade.sh does the whole sequence — snapshot, checkout, npm ci --omit=dev, restart, and
report the running version. It defaults to the newest stable tag, deliberately skipping
-rc/-beta/-alpha prereleases:
cd /opt/screentinker
scripts/upgrade.sh # latest stable release
scripts/upgrade.sh v1.2.3 # or pin one
If you are doing it by hand, the order matters:
sudo -u <service-user> git fetch --tags origin
sudo -u <service-user> git checkout -f v1.2.3
# only if dependencies actually changed:
cd server && sudo -u <service-user> npm ci --omit=dev
sudo systemctl restart <service>
Ownership first. Every file must belong to the service user before the checkout. A checkout
that fails partway through leaves the worst possible state: VERSION updated while the code is
still the old release, so the service reports a version it is not running and no migrations ran.
sudo chown -R <service-user>:<service-user> /opt/screentinker
A service user with no home directory breaks npm. It writes logs and a cache to $HOME, which
does not exist, and installs nothing while looking like it worked:
cd server && sudo -u <service-user> env HOME=/opt/screentinker \
npm_config_cache=/opt/screentinker/.npm-cache npm ci --omit=dev
Prove the checkout is complete — a version string alone will not tell you:
git status --porcelain --untracked-files=no # want: empty
git diff <tag> -- server frontend # want: empty
Deploying: Docker
# in the compose directory
cp -a docker-compose.yml docker-compose.yml.bak-pre-<version>
sed -i 's|screentinker:<old>|screentinker:<new>|' docker-compose.yml
docker compose pull && docker compose up -d
Migrations run at boot exactly as they do natively. State lives in the named volume (st-data in
the example compose), so recreating the container does not touch the database.
Anything bind-mounted into the container — the served APK, a .wgt, custom assets — must be updated
on the host, and see the inode warning below.
The served APK
The file the OTA endpoint hands to Android displays. Two rules, both learned the hard way.
1. Replace it in place. Never mv or cp over it.
It is a bind-mounted file, so the container holds the inode. Replacing the file gives the host a new inode and the container keeps serving the old bytes forever, with nothing in any log to say so.
cat /tmp/new.apk > /opt/screentinker/ScreenTinker.apk # correct — same inode
# NOT: mv, cp, install, or anything that unlinks and recreates
stat -c %i /opt/screentinker/ScreenTinker.apk # confirm it did not change
2. The advertised size must match the served bytes, or displays loop.
/api/update/check reports apk_size from a cache refreshed every OTA_APK_REFRESH_MS
(default 60s), and the server re-stats at boot. If the advertised size and the real file disagree,
a display downloads, rejects, and retries — forever. After swapping, restart the service and confirm:
curl -s 'http://127.0.0.1:3001/api/update/check?version=<an-older-version>'
stat -c %s /opt/screentinker/ScreenTinker.apk # must equal the reported apk_size
⚠️ The query parameter is
version, notcurrent_version. The wrong name yieldsreason: no-version, update_available: false, which looks exactly like a broken OTA but is not./api/versionis a different endpoint and itsupdate_availableis not the OTA verdict.
Verify the signature after any APK swap, and use jarsigner:
jarsigner -verify ScreenTinker.apk # want: "jar verified."
unzip -l ScreenTinker.apk | grep META-INF # want: a .SF and a .RSA
apksigner verify -v misreports v1 scheme: false on some build-tools versions even when the JAR
signature is present and valid. MDM-managed signage needs v1, so trust jarsigner.
Verifying a deploy
curl -s http://127.0.0.1:3001/api/version # version + build hash
curl -s http://127.0.0.1:3001/api/status # health, loop lag, connected displays
Then, and this is the part people skip:
- Row counts match what you recorded beforehand.
- The log is clean. Migrations reported, no errors:
docker logs <container> 2>&1 | grep -iE 'migrat|error|exception' # or journalctl -u <service> - Check through your reverse proxy / CDN too, not only on loopback. Cached or misrouted assets only show up from outside.
⚠️ A version string is not proof the new code is running, and neither is the build hash. The hash covers the frontend, so a server-only change deploys with an unchanged hash — which looks exactly like a stale image. When it matters, check for the code itself:
docker exec <container> grep -c '<a symbol only the new version has>' /app/server/<file>
A frontend change needs a hard refresh (Ctrl+Shift+R) before you judge it. Assets revalidate, but a browser sitting on the old bundle will show you the old behaviour and you will debug a fixed bug.
Rolling back
Because backups are taken per deploy, rollback is mechanical:
Native
sudo -u <service-user> git checkout -f <previous-tag>
cd server && npm ci --omit=dev # only if dependencies changed
cat /path/to/ScreenTinker.apk.bak > /opt/screentinker/ScreenTinker.apk
sudo systemctl restart <service>
Docker
cp -a docker-compose.yml.bak-<version> docker-compose.yml
cat /path/to/ScreenTinker.apk.bak > /opt/screentinker/ScreenTinker.apk
docker compose up -d
The database usually does not need restoring. Migrations are additive, so older code simply ignores the new columns. Restore the snapshot only if a migration was destructive — and if one ever is, that is the moment to stop and read it rather than reflexively rolling forward.
Releases and version numbers
Cutting a release is documented in RELEASING.md. The operational consequences:
A prerelease sorts BELOW its own release. 1.2.3-alpha1 is semver-older than 1.2.3. That has
two effects worth internalising:
- A display that takes a prerelease is not "ahead"; a later stable of the same version supersedes it, which is what you want.
- The Android update check offers a prerelease to any older client on the stable channel. Putting a prerelease on an instance means every Android display below it takes it at its next check. Do that deliberately, on an instance whose displays you are willing to move.
:latest is not moved for a prerelease. The release workflow skips it for any tag containing a
-, so nobody tracking :latest pulls untested code on their next restart.
Android versionCode must never go backwards. Android refuses a downgrade, so a build with a
lower code cannot install over a higher one — the usual cause is a side-loaded test build whose code
was bumped past the release line. Keep the release line ahead of anything you side-load, or you will
be reinstalling by hand (which wipes app data and drops pairing).
A re-cut tag is only safe if it published nothing. If a tag has already produced a GitHub Release or an image, delete-and-repush is not a fix; cut the next version instead.
Traps worth knowing before they bite
Native modules are built for one Node ABI. better-sqlite3 is compiled against the Node that
installed it. Run the app — or its tests — under a different major version and it fails with
NODE_MODULE_VERSION mismatch, which presents as hundreds of unrelated test failures rather than
one clear error. Use the same Node the service runs.
SQLite foreign keys are off unless enabled per connection. A declared ON DELETE CASCADE does
not fire on its own, so deleting a parent row can leave orphaned children. Check with
PRAGMA foreign_key_check; after any bulk delete.
Deploying reloads every connected web player. The frontend self-reloads when the build hash changes. Browsers cope. Some embedded webview players do not, and may need a restart afterwards — worth knowing before you deploy during business hours.
An SSO-linked administrator has no password. If you link the platform administrator account to
an identity provider and that provider later fails, the login page cannot help you. Recovery is
node scripts/reset-admin.js on the server. See sso-setup.md.
Backups are only real once restored. A snapshot that has never been restored is a hypothesis.
Periodically restore the newest one into a throwaway instance and confirm it boots and serves
/api/status.