Symptom 1's 'stuck on first load, fixed by toggling the playlist' is stale
download backoff. On a fresh device the first downloads fail while the link is
settling; DownloadCoordinator escalates an exponential backoff (15s..5min cap),
and ensure() then SKIPS those items. The 60s playlist refresh re-fires
onPlaylistUpdate but ensure() still skips them, and backoff was only ever cleared
by forget() (content-delete) — never by a re-assignment. So the item stays stuck
until the 5-min window happens to lapse; toggling the playlist is just a manual
way to wait it out.
Fix (storm-safe — neither reset fires on the routine same-playlist 60s refresh):
1. DownloadCoordinator.resetBackoff(id) / resetAllBackoff() — clear attempts +
nextAttemptAt but KEEP inFlight (single-flight preserved, no duplicate .part).
2. onPlaylistUpdate resets backoff for each item ONLY when the content-id
signature changed (first load, reassignment, toggle-back), then ensures — so a
genuine (re)assignment retries immediately. Same-signature 60s refresh -> no
reset -> the retry-storm guard stays intact.
3. Network onAvailable (was onLost-only) -> resetAllBackoff() + requestPlaylistRefresh,
so content that failed while the link was settling retries the moment real
connectivity arrives.
Tests: DownloadCoordinatorTest — resetBackoff/resetAllBackoff re-attempt a
backed-off item before the clock advances; existing backoff/single-flight tests
unchanged. :app:testDebugUnitTest green.
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
The player opened duplicate/rapid WebSocket connections for the same device_id: connect() was
unconditional (disconnect + forceNew socket) and reachable from every lifecycle entry point
(boot, service start, MainActivity/ProvisioningActivity bind, foreground re-bind, START_STICKY).
A ROM that re-binds on foreground (MAXHUB PROC_STATE_TOP, isBindService:true) therefore
re-invoked connect() repeatedly -> a burst of sockets, each evicted by the next (the 8-in-9s
storm). Fire TV never re-binds like that, so it never reproduced.
- ConnectionGuard (new, pure/testable — service is the shell, per the OtaThrottle pattern):
shouldOpenNewSocket(hasSocket, sameUrl, socketActive) — reuse a live/self-healing socket to
the same url; open a new one only when none is usable.
- WebSocketService: connect() is now idempotent (@Synchronized + ConnectionGuard) — every entry
point reuses the one socket, never opens a duplicate; body split into openSocket(). socketActive
/ currentUrl track the single socket.
- Single owner: onStartCommand now calls connect() so the SERVICE owns the one connection
(idempotent across START_STICKY restarts), not whichever activity binds.
- Reconnect discipline: on io server/client disconnect (which Socket.IO does NOT auto-reconnect)
mark the socket inert and schedule exactly ONE backed-off re-open — never a blind re-open loop;
a transport drop keeps socketActive=true so Socket.IO's own reconnect is reused.
Test: ConnectionGuardTest (5, incl. 8-rapid-binds-all-reuse). :app:testDebugUnitTest green
(ConnectionGuard 5, OtaThrottle 7, ScheduleEval 1). NOT bumped/signed/released — Dan builds+signs
with the BMG keystore; 1.9.2-patch2 (server net) covers un-updated devices.
Follow-up to the cache/backoff loop fix (aa23cf0): make a device that can't
self-install visible to operators, and fix the signature-verify bug that kept the
whole #139 fix from engaging on the actual Fire OS target.
Dashboard surface (Phase 2):
- devices gains ota_status / ota_target_version / ota_attempts / ota_updated_at
via the idempotent ALTER TABLE ADD COLUMN migration (non-destructive,
default-backfilled, idempotent on re-run).
- The device reports ota_status (OtaThrottle.statusFor -> none | pending |
manual_update_required) in device_info; the server persists it on register
(the reconnect backstop). devices d.* already surfaces it to the dashboard.
- Dashboard shows a non-blocking amber badge when manual_update_required
("Update available (vX) - install failed N times, manual update required");
i18n key in en.js (non-en inherits via the en fallback). Server suite +1 test.
Event-driven status (Option B):
- New device:ota-status WS message, emitted on STATE TRANSITIONS only
(enter-backoff -> manual_update_required, clear -> none), so the badge updates
promptly without waiting for a reconnect and without per-poll/heartbeat chatter.
Server handler persists the same fields; an unknown/forged device_id is a safe
no-op. The register-path persist stays as the reconnect backstop.
Signature-verify fix (the critical piece):
verifyApkSignature read the downloaded APK's signer via
getPackageArchiveInfo(GET_SIGNING_CERTIFICATES).signingInfo, but that field is
null for ARCHIVE files on API 28/29 (populated only from API 30). On Fire OS 8
(Android 9 / API 28) - the actual deployment target - this returned 0 certs from
a correctly-signed APK, so every OTA was refused as "tampered," the cache was
deleted, and the full APK re-downloaded every check cycle. This was the real
cause of the #139 re-download loop, NOT a silent-install failure: the cache and
backoff added in this branch sit behind this verify gate and never engaged on
the target.
Fix: below API 30, read the archive's signer via the legacy GET_SIGNATURES +
.signatures (its v1/JAR cert, which IS populated on 28/29). Keep
GET_SIGNING_CERTIFICATES + signingInfo for API >= 30 and for the installed-app
read (which works on 28+). The archive's signer is still extracted and compared
to the installed app's signer; a mismatch or zero-cert APK is still rejected.
This reads the cert correctly on old APIs - it does not weaken verification.
Verified on emulators:
- API 28: verify now passes for a legit APK (was: 0 certs, refused). Full backoff
then engages - 8.5MB pulled once, cache-hit on retries, backoff after 3,
manual_update_required emitted once; clears on successful update.
- API 28 negative: a re-signed (different-key) APK is still refused on cert
MISMATCH - no hole opened.
- API 30: unchanged path still passes (no regression).
- server suite 173/173, OtaThrottleTest 7/7.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Devices that download an OTA APK but cannot silently install it (Fire TV: no
device-owner path) re-downloaded the full APK every check cycle indefinitely -
install never completes, version never advances, next check re-triggers.
Client (UpdateChecker.kt, ServerConfig.kt, OtaThrottle.kt):
- Reuse a cached, signature-verified APK instead of re-downloading every cycle;
delete leftover invalid files; keep the verified APK on disk as the
manual-install artifact.
- Persisted per-version attempt budget (EncryptedSharedPreferences) so it
survives the Fire OS app restarts that drive the loop. An attempt is counted
only when an install is launched - a download/verify failure does not consume
the budget, so a transient network problem cannot park a healthy device in
backoff. After 3 failed installs, back off to one retry per 24h.
- Clear OTA state and caches when a check returns update_available=false while
state is pending (app relaunched as the new version).
- Report OTA status to the dashboard via device:log (tag ota) on state
transitions only (enter-backoff, clear) to avoid flooding the channel.
- Extract throttle decision logic into a pure OtaThrottle object (no Android
deps) with JUnit coverage (OtaThrottleTest) for the state transitions.
Server (server.js):
- Reword /download/apk log from "OTA update in progress" to "APK served" and
rate-limit to once per IP / 10 min so a looping device cannot flood the log.
Note: client-cooperative fix - prevents the loop in cohorts running this APK.
Currently-stuck beta4 devices still require a one-time manual update.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Each playlist item can carry schedule blocks (active days, start/end
time-of-day, optional start/end dates). An item plays when the screen's
local "now" matches at least one block; an item with no blocks always
plays. #74 covers time-of-day/day-of-week windows including overnight
wrap; #75 covers inclusive date ranges (auto-expiry). Evaluation is
on-device, so dayparting and expiry work offline.
- Shared evaluator contract: shared/schedule-vectors.json (39 vectors —
DST US+AU, overnight-wrap anchoring, timezone correctness, date
boundaries). Canonical JS evaluator in server/lib/schedule-eval.js;
Kotlin and Tizen ports kept in lockstep by drift guards (Tizen byte-diff
test, Kotlin JUnit reads the shared JSON, new android-test CI job).
- All three players (web, Android, Tizen) filter by schedule against their
own clock, idle with a "Nothing scheduled" message + 30s re-check when
everything is filtered, and fail open on any evaluator error.
- Editor: per-item schedule modal + row badge in the playlist editor;
client validation mirrors the server; editing marks the playlist draft.
- Part B (behaviour change): device/group schedule overrides now evaluate
in each device's effective timezone instead of server-local time.
- Device detail shows the reported timezone + a clock-skew warning.
- i18n for en/es/fr/de/pt across all new strings (namespaced itemsched.*
to avoid colliding with the device-schedule calendar's schedule.*).
- CHANGELOG documents the feature, the Part B change, the fail-open
guarantee, and the scheduled-single-video re-render tradeoff.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>