mirror of
https://github.com/screentinker/screentinker.git
synced 2026-08-13 13:53:12 -06:00
* spike: replace sharp with pure-JS image ops (jimp + jsquash WASM) Removes the last native dependency from the ingest path, so the server no longer needs a per-platform/per-ABI prebuilt to thumbnail an image. Motivated by getting the server onto hardware with no toolchain, but the ABI tax is paid on every install — it is the same failure class lib/preflight-deps.js exists to explain. lib/image-ops.js is the whole surface: metadata() and writeThumbnail(), which are the only two things ingest ever asked sharp for. Format parity holds. jpeg/png/gif/tiff/bmp are native to Jimp; webp and avif go through @jsquash WASM, whose bundled .wasm must be compiled by hand because the packages locate it with fetch(file://) and Node has no file:// fetch — the only symptom otherwise is a bare "fetch failed". heic is unsupported, as it already was: sharp advertises heif but its prebuilt libvips refuses HEVC. #170 is preserved by a different mechanism. Jimp applies EXIF orientation at decode and rewrites the tag to 1, so metadata() reports display dimensions and imageDisplayDims() runs as a no-op instead of swapping W/H a second time. The helper stays in the path so the rule keeps living in one place. Verified: 1643/1643 tests pass, and ingest was exercised in a child process with node_modules/sharp moved aside — jpeg, EXIF-rotated jpeg, png, webp, avif, gif all measured and thumbnailed correctly, corrupt input still yields nulls with no phantom thumbnail_path. KNOWN BLOCKER, do not ship as-is: Jimp is pure JS on the main thread, where sharp handed work to a libvips threadpool. A 12MP photo goes 65ms -> 1079ms, and the event loop stalls for 1003ms of it (sharp: zero stalls). thumbnail-backfill.js walks a whole library at boot, so this reproduces #240 exactly — blocked loop, missed heartbeats, panels marked offline, reconnect churn. Needs a worker_thread offload before this is viable; image-ops.js is the seam for it. * Run image decoding on a worker thread Fixes the blocker the previous commit shipped with. Pure-JS decoding costs ~1s of solid CPU for a 12MP photo, and in-process that is not a slow upload but a stalled event loop — no heartbeats, no socket traffic. thumbnail-backfill.js walks a whole library at boot, so it reproduced #240 (blocked loop -> missed heartbeats -> panels offline -> reconnect churn) from our own maintenance. sharp never did this because libvips works on a threadpool. image-ops.js is now a dispatcher over image-ops-worker.js; the work moved unchanged to image-ops-core.js, so callers and their failure contract are untouched. Measured on a 12MP photo: 1079ms wall with the loop stalled 1003ms, to 1881ms wall for two ops with ZERO stalls and 185 timer ticks serviced. Wall time is worse and that is fine — it is off the main thread now. Design notes, all load-bearing: - ONE JOB AT A TIME. A decoded 12MP bitmap is ~48MB of RGBA; overlapping jobs multiply peak memory by queue depth, which is the wrong failure on the small targets this change exists to reach. Costs no throughput — the work is CPU-bound and one busy worker already saturates its core. - unref'd while idle, ref'd only in flight. Otherwise scripts/backfill-rotation- dims.js never exits and `node --test` hangs forever. Verified: a CLI-style run exits in 104ms, code 0. - decode failures reply as messages, so one bad upload cannot tear down the worker and take unrelated queued jobs with it. - in-process fallback if a thread cannot be had, warned rather than silent. test/image-ops.test.js pins the loop-liveness property, which no functional test would catch. Its thresholds were mutation-tested against the inline path: the first version passed there too (4MP stalls only ~355ms, under a non-flaky threshold), so the fixture is 12MP and the thresholds sit in the gap between the two behaviours — worker ~90 ticks/~0ms, inline ~3 ticks/~897ms. It now fails inline, as a guard must. 1647/1647 pass. Ingest re-verified with node_modules/sharp moved aside. * Measure and thumbnail an image from a single decode Ingest asked for metadata() then writeThumbnail(), which decoded the file twice. That pairing was free under sharp, whose .metadata() only parses the header, but every decode here is a full one — ~1s for a 12MP photo — so the naive translation doubled the most expensive thing on the ingest path. image-ops.measureAndThumbnail() returns both from one decode. Full ingest of a 12MP photo: 2 decodes/~1.9s -> 1150ms, still with zero event-loop stalls. The subtlety is the failure contract. In the two-call version width and height were assigned BEFORE the thumbnail was attempted, so a failed thumbnail still left usable dimensions on the row — the player needs them to letterbox. Merging naively would have turned any thumbnail failure into total metadata loss. So a WRITE failure is reported ({thumbnailWritten:false, thumbnailError}) with the dimensions intact, and the caller sets thumbnail_path only when the write succeeded, keeping the phantom-path discipline. A DECODE failure still throws — there is nothing to report about an unreadable image. backfill-rotation-dims.js deliberately keeps the separate calls: it probes every image row but regenerates a thumbnail only for the few whose dimensions changed, so pairing them there would decode files it has no reason to thumbnail. Tests count decodes rather than timing them — an exact property, and a wall-clock comparison would be flaky under load. The count filters for reads of the file under test: Node's ESM loader also goes through fs.promises.readFile, so a raw call count picks up jimp's and the WASM codecs' lazy loading and reads 30 instead of 1. 1649/1649 pass. Ingest re-verified across all 7 formats with sharp moved aside. * Dockerfile: sharp is no longer a production dependency --omit=dev now leaves it out entirely; better-sqlite3 is the only native module the builder stage still needs a toolchain for.
58 lines
3 KiB
Docker
58 lines
3 KiB
Docker
# ScreenTinker server image: serves the dashboard, the web player, and the
|
|
# device API. All mutable state (db, uploads, jwt secret) lives under /data so it
|
|
# survives container restarts - mount a volume there. A built ScreenTinker.apk
|
|
# can be mounted at /data/ScreenTinker.apk to enable OTA APK downloads.
|
|
#
|
|
# No TLS in the image: it listens on plain HTTP :3001. Front it with a
|
|
# TLS-terminating reverse proxy / Cloudflare in production.
|
|
|
|
# --- builder: install production deps (better-sqlite3 is the only native one left; image
|
|
# decoding is pure JS + WASM since sharp was dropped, and sharp is now a devDependency that
|
|
# --omit=dev leaves out entirely) ---
|
|
FROM node:20-slim AS builder
|
|
WORKDIR /app/server
|
|
# build toolchain in case a native prebuild is missing for the target arch
|
|
RUN apt-get update \
|
|
&& apt-get install -y --no-install-recommends python3 build-essential \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
COPY server/package.json server/package-lock.json ./
|
|
RUN npm ci --omit=dev
|
|
|
|
# --- runtime ---
|
|
FROM node:20-slim
|
|
# ffmpeg (ships ffprobe) powers video thumbnails + duration extraction at upload.
|
|
# Without it videos still upload and play, but arrive with no thumbnail or duration.
|
|
RUN apt-get update \
|
|
&& apt-get install -y --no-install-recommends ffmpeg \
|
|
&& rm -rf /var/lib/apt/lists/*
|
|
ENV NODE_ENV=production
|
|
# Relocate all state onto the volume (config.js reads DATA_DIR; unset would use
|
|
# the in-repo paths, which we do not want in a container).
|
|
ENV DATA_DIR=/data
|
|
WORKDIR /app/server
|
|
# App source (node_modules/test/db/uploads/certs are excluded via .dockerignore),
|
|
# then the built deps, the frontend the server serves, and the VERSION file it
|
|
# reads as ../VERSION.
|
|
COPY server/ /app/server/
|
|
COPY --from=builder /app/server/node_modules /app/server/node_modules
|
|
COPY frontend/ /app/frontend/
|
|
# shared/Transitions is a RUNTIME dependency: server/lib/transition-config.js + transition-bundle.js
|
|
# require the shader manifest/params/sources from ../../shared at load time (the server won't boot
|
|
# without it). Small, and keeps the .glsl files the single source across server + player + Tizen.
|
|
COPY shared/ /app/shared/
|
|
COPY VERSION /app/VERSION
|
|
# the /openapi.yaml route serves ../docs/openapi.yaml (the spec Redoc on /docs fetches);
|
|
# without this it 404s in the image even though it serves fine from a dev checkout.
|
|
COPY docs/openapi.yaml /app/docs/openapi.yaml
|
|
# database.js requires scripts/migrate-multitenancy at boot
|
|
COPY scripts/ /app/scripts/
|
|
# The BrightSign bridge and sync modules are served to the player from ../brightsign so the copy
|
|
# the player loads can never drift from the one on the player's own storage. That RUNTIME path
|
|
# does not exist unless the directory is in the image: without this the routes 404 in a container
|
|
# while working perfectly from a dev checkout — and a missing player asset fails silently, because
|
|
# the SPA fallback answers 200 with HTML where JavaScript was expected.
|
|
COPY brightsign/ /app/brightsign/
|
|
VOLUME ["/data"]
|
|
EXPOSE 3001
|
|
CMD ["node", "server.js"]
|