The firmware publishes status/heartbeat, system/alerts, system/info and
status/playback with retain=true. On every backend (re)connect - every
restart and every uvicorn --reload - the broker replays the last message on
each of those topics for every device that ever connected. We handled those
replays as if they had just happened:
- heartbeats: a row with received_at=now() for every device, so devices
that have been dead for months showed ONLINE for 90s after each restart
and got pinged. ~770k such rows exist locally.
- boot_report: the last boot logged again as a new reboot (the phantom
PANIC entries on the Health tab).
- alerts / other info events: logged again as new occurrences.
MQTT delivers retain=1 only for replays caused by a new subscription; live
publishes always arrive with retain=0. The flag is now passed through to the
handlers and the WS broadcast:
- heartbeat: replays are not stored. A live heartbeat is.
- {"state":"offline"} heartbeat (LWT / graceful disconnect) is no longer
stored as a sign of life. It marks the device offline immediately in
a small in-memory set (mqtt/presence.py) used by /mqtt/status and the ping
loop; a later live heartbeat clears it. Replayed offline markers also mark
offline, since a retained message is the device's last word.
- boot_report: live -> always a new boot. Replay -> stored only if it
differs from the device's latest boot row (i.e. we missed it while down).
- alerts: replay still syncs the current-alert row; history gets a row on a
live alert (even an identical repeat - faults recur) or on a replay that
changes state. Replaces the 98dd16b rule that dropped identical live alerts.
- other info events: replays are not logged.
- Frontend (DeviceList, DeviceDetail, LogsTab) ignores retained WS messages
for live updates, and flips a device offline on the offline marker instead
of marking it online.
Verified locally after a backend restart: only the 7 actually-live devices
got new heartbeat rows (none from the replays), and no boot/alert/info rows
were created.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The firmware publishes boot_report on system/info and alerts on
system/alerts with retain=true. Every time the backend (re)connects - which
under uvicorn --reload means every backend file save - the broker redelivers
the last retained message and we inserted it again with occurred_at=now().
Result: the Health tab showed fresh PANIC boots and "Device reset due to
fault" alerts for a device that had been up for 4 days.
- insert_boot_event skips the insert when a row with the same
(device_serial, boot_count) already exists. boot_count is the firmware's
lifetime counter, so it uniquely identifies a boot.
- upsert_alert only writes when state/message actually changed and returns
whether it did; the alert-event history row is only added on a change.
A redelivered identical alert no longer bumps updated_at either.
Existing duplicate rows are not touched by this commit.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The shared legacy password is still needed for boards on pre-HMAC
firmware, but it was accepted for any username. Now:
- controlled by MQTT_ALLOW_LEGACY_PASSWORD (config.py, default true;
documented in .env.example) so it can be switched off without a deploy,
- only accepted for device-shaped usernames (uppercase alphanumeric
segments joined by "-", optional "-kiosk"), never for app_ users or
any other shape,
- every successful legacy login is logged at WARNING with the username,
rate-limited to once per username per hour, so the boards still
depending on it are visible before the flag is turned off.
HMAC auth is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
POST /mqtt/auth/acl now handles "app_<uid>" users:
- topic must be exactly vesper/{serial}/<a>/<b> with serial in the user's
device_serials (resolved by the users doc `uid` field),
- publish (acc 2): only control/command,
- subscribe (acc 4) and read/delivery (acc 1): only control/ack,
status/heartbeat, status/playback,
- wildcard topics (+ / #) are denied,
- clientid must start with "app_<uid>_" so one user cannot reuse another
user's client id to kick them off,
- blocked users are denied; anything else (incl. other acc values) is 403.
Lookups go through mqtt/app_users.py's 60s TTL cache so per-message
checks don't hit Firestore every time; assign/unassign/block invalidate it.
Also fixes the acc comment: mosquitto passes 1 = read (delivery),
2 = write (publish), 4 = subscribe - not "1 = subscribe, 3 = both".
Device/kiosk ACL is unchanged.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The remote FlutterFlow app connects to Mosquitto as "app_<firebase_uid>"
with a Firebase ID token as the password, so no per-user MQTT accounts
need to exist anywhere.
For app_ usernames, POST /mqtt/auth/user now:
- verifies the token with firebase_admin.auth.verify_id_token
(check_revoked=True),
- requires the decoded uid to equal the uid in the username,
- requires a users doc with that `uid` field (queried, not by doc id)
whose status is not "blocked" (same meaning as users.service.block_user).
It returns 200/403 and logs the deny reason - never the token.
app_ usernames never fall through to the HMAC / legacy "vesper" check.
Device and kiosk auth are unchanged. App users are still denied every
topic by the existing ACL until the app ACL lands in the next commit.
Both handlers are now plain `def` so the blocking Firestore / Firebase
calls run in FastAPI's threadpool instead of stalling the event loop.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds a `device_serials: [string]` array to Firestore `users` docs so the
MQTT ACL (and get_user_devices) can answer "which boards may this user
reach?" without streaming the entire devices collection.
- The serial is the value used in MQTT topics vesper/{serial}/...: the
device doc's `serial_number` (flashed into NVS, used by the firmware as
its MQTT id), falling back to the legacy `device_id` for old docs.
Centralised in users.service.device_serial_of().
- assign_device / unassign_device now write the device's user_list and the
user's device_serials (ArrayUnion/ArrayRemove) in one atomic batch.
- The device Manage tab endpoints (POST/DELETE /api/devices/{id}/user-list)
also edit user_list, so they get the same batched sync - otherwise the
most common assignment path would silently leave device_serials stale.
- get_user_devices resolves devices via device_serials with chunked
Firestore "in" queries instead of a full collection scan. Requires the
backfill script (next commit) to be run for existing assignments.
- New mqtt/app_users.py: resolves users by the `uid` FIELD (not doc id -
create_user uses .add(), FlutterFlow uses uid as doc id) with a 60s
in-process TTL cache. Assign/unassign, update, block/unblock and delete
invalidate that uid's entry.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two efforts that landed together because the v2 topic work extends
tables the health-telemetry effort added days earlier in the same
files/functions, making them impractical to separate cleanly:
Health/diagnostics telemetry (schema, Jul 13-17):
- New Postgres tables: device_alert_events, device_boot_events,
device_ping_samples, device_diagnostics_reports, plus a `source`
column on device_logs to distinguish log origins
- Query/service layer in pg_mqtt.py and database/__init__.py for
inserting and listing this history, plus a "latest metrics" endpoint
combining most-recent diagnostics + ping RTT per device
- mqtt/router.py gains list endpoints for alert/boot/ping/diagnostics
history, consumed by the upcoming Health tab
MQTT v2 topic migration (Sep 21):
- Heartbeat payload flattened per vesper_mqtt_topic_spec_v2.md, adding
rssi/free_heap/state/ok fields
- Command replies move to control/ack, device-initiated events to
control/reports; mqtt/client.py subscribes to the new topic set and
runs a ping_loop (wired up in main.py) for RTT sampling
- mqtt/logger.py and pg_mqtt.py updated to parse and persist the new
payload shape alongside the legacy fields for backwards compatibility
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>