mirror of
https://github.com/screentinker/screentinker.git
synced 2026-08-13 22:03:13 -06:00
* spike: replace sharp with pure-JS image ops (jimp + jsquash WASM) Removes the last native dependency from the ingest path, so the server no longer needs a per-platform/per-ABI prebuilt to thumbnail an image. Motivated by getting the server onto hardware with no toolchain, but the ABI tax is paid on every install — it is the same failure class lib/preflight-deps.js exists to explain. lib/image-ops.js is the whole surface: metadata() and writeThumbnail(), which are the only two things ingest ever asked sharp for. Format parity holds. jpeg/png/gif/tiff/bmp are native to Jimp; webp and avif go through @jsquash WASM, whose bundled .wasm must be compiled by hand because the packages locate it with fetch(file://) and Node has no file:// fetch — the only symptom otherwise is a bare "fetch failed". heic is unsupported, as it already was: sharp advertises heif but its prebuilt libvips refuses HEVC. #170 is preserved by a different mechanism. Jimp applies EXIF orientation at decode and rewrites the tag to 1, so metadata() reports display dimensions and imageDisplayDims() runs as a no-op instead of swapping W/H a second time. The helper stays in the path so the rule keeps living in one place. Verified: 1643/1643 tests pass, and ingest was exercised in a child process with node_modules/sharp moved aside — jpeg, EXIF-rotated jpeg, png, webp, avif, gif all measured and thumbnailed correctly, corrupt input still yields nulls with no phantom thumbnail_path. KNOWN BLOCKER, do not ship as-is: Jimp is pure JS on the main thread, where sharp handed work to a libvips threadpool. A 12MP photo goes 65ms -> 1079ms, and the event loop stalls for 1003ms of it (sharp: zero stalls). thumbnail-backfill.js walks a whole library at boot, so this reproduces #240 exactly — blocked loop, missed heartbeats, panels marked offline, reconnect churn. Needs a worker_thread offload before this is viable; image-ops.js is the seam for it. * Run image decoding on a worker thread Fixes the blocker the previous commit shipped with. Pure-JS decoding costs ~1s of solid CPU for a 12MP photo, and in-process that is not a slow upload but a stalled event loop — no heartbeats, no socket traffic. thumbnail-backfill.js walks a whole library at boot, so it reproduced #240 (blocked loop -> missed heartbeats -> panels offline -> reconnect churn) from our own maintenance. sharp never did this because libvips works on a threadpool. image-ops.js is now a dispatcher over image-ops-worker.js; the work moved unchanged to image-ops-core.js, so callers and their failure contract are untouched. Measured on a 12MP photo: 1079ms wall with the loop stalled 1003ms, to 1881ms wall for two ops with ZERO stalls and 185 timer ticks serviced. Wall time is worse and that is fine — it is off the main thread now. Design notes, all load-bearing: - ONE JOB AT A TIME. A decoded 12MP bitmap is ~48MB of RGBA; overlapping jobs multiply peak memory by queue depth, which is the wrong failure on the small targets this change exists to reach. Costs no throughput — the work is CPU-bound and one busy worker already saturates its core. - unref'd while idle, ref'd only in flight. Otherwise scripts/backfill-rotation- dims.js never exits and `node --test` hangs forever. Verified: a CLI-style run exits in 104ms, code 0. - decode failures reply as messages, so one bad upload cannot tear down the worker and take unrelated queued jobs with it. - in-process fallback if a thread cannot be had, warned rather than silent. test/image-ops.test.js pins the loop-liveness property, which no functional test would catch. Its thresholds were mutation-tested against the inline path: the first version passed there too (4MP stalls only ~355ms, under a non-flaky threshold), so the fixture is 12MP and the thresholds sit in the gap between the two behaviours — worker ~90 ticks/~0ms, inline ~3 ticks/~897ms. It now fails inline, as a guard must. 1647/1647 pass. Ingest re-verified with node_modules/sharp moved aside. * Measure and thumbnail an image from a single decode Ingest asked for metadata() then writeThumbnail(), which decoded the file twice. That pairing was free under sharp, whose .metadata() only parses the header, but every decode here is a full one — ~1s for a 12MP photo — so the naive translation doubled the most expensive thing on the ingest path. image-ops.measureAndThumbnail() returns both from one decode. Full ingest of a 12MP photo: 2 decodes/~1.9s -> 1150ms, still with zero event-loop stalls. The subtlety is the failure contract. In the two-call version width and height were assigned BEFORE the thumbnail was attempted, so a failed thumbnail still left usable dimensions on the row — the player needs them to letterbox. Merging naively would have turned any thumbnail failure into total metadata loss. So a WRITE failure is reported ({thumbnailWritten:false, thumbnailError}) with the dimensions intact, and the caller sets thumbnail_path only when the write succeeded, keeping the phantom-path discipline. A DECODE failure still throws — there is nothing to report about an unreadable image. backfill-rotation-dims.js deliberately keeps the separate calls: it probes every image row but regenerates a thumbnail only for the few whose dimensions changed, so pairing them there would decode files it has no reason to thumbnail. Tests count decodes rather than timing them — an exact property, and a wall-clock comparison would be flaky under load. The count filters for reads of the file under test: Node's ESM loader also goes through fs.promises.readFile, so a raw call count picks up jimp's and the WASM codecs' lazy loading and reads 30 instead of 1. 1649/1649 pass. Ingest re-verified across all 7 formats with sharp moved aside. * Dockerfile: sharp is no longer a production dependency --omit=dev now leaves it out entirely; better-sqlite3 is the only native module the builder stage still needs a toolchain for. |
||
|---|---|---|
| .. | ||
| backfill-rotation-dims.js | ||
| demo-schedule.js | ||