fix(mqtt): treat broker-replayed retained messages as stale state

The firmware publishes status/heartbeat, system/alerts, system/info and
status/playback with retain=true. On every backend (re)connect - every
restart and every uvicorn --reload - the broker replays the last message on
each of those topics for every device that ever connected. We handled those
replays as if they had just happened:

- heartbeats: a row with received_at=now() for every device, so devices
  that have been dead for months showed ONLINE for 90s after each restart
  and got pinged. ~770k such rows exist locally.
- boot_report: the last boot logged again as a new reboot (the phantom
  PANIC entries on the Health tab).
- alerts / other info events: logged again as new occurrences.

MQTT delivers retain=1 only for replays caused by a new subscription; live
publishes always arrive with retain=0. The flag is now passed through to the
handlers and the WS broadcast:

- heartbeat: replays are not stored. A live heartbeat is.
- {"state":"offline"} heartbeat (LWT / graceful disconnect) is no longer
  stored as a sign of life. It marks the device offline immediately in
  a small in-memory set (mqtt/presence.py) used by /mqtt/status and the ping
  loop; a later live heartbeat clears it. Replayed offline markers also mark
  offline, since a retained message is the device's last word.
- boot_report: live -> always a new boot. Replay -> stored only if it
  differs from the device's latest boot row (i.e. we missed it while down).
- alerts: replay still syncs the current-alert row; history gets a row on a
  live alert (even an identical repeat - faults recur) or on a replay that
  changes state. Replaces the 98dd16b rule that dropped identical live alerts.
- other info events: replays are not logged.
- Frontend (DeviceList, DeviceDetail, LogsTab) ignores retained WS messages
  for live updates, and flips a device offline on the offline marker instead
  of marking it online.

Verified locally after a backend restart: only the 7 actually-live devices
got new heartbeat rows (none from the replays), and no boot/alert/info rows
were created.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-30 16:01:41 +03:00
co-authored by Claude Opus 5.5
parent 18e57c6e5f
commit 46c5c0a846
8 changed files with 111 additions and 27 deletions
+7 -10
View File
@@ -444,20 +444,17 @@ async def insert_boot_event(device_serial: str, boot_count: int | None,
crash_task: str | None = None,
crash_pc: int | None = None,
crash_exc_cause: int | None = None,
crash_exc_vaddr: int | None = None) -> int | None:
"""Insert a boot event, unless it is a repeat of the device's most recent
one (returns None then). boot_report is published retained on system/info,
so the broker redelivers the last one every time the backend (re)subscribes
— without this check each backend restart logged the device's last boot
(crash detail included) again as a brand-new event.
crash_exc_vaddr: int | None = None,
skip_if_latest: bool = False) -> int | None:
"""Insert a boot event. With skip_if_latest=True (a retained replay of the
device's last boot_report), skip it and return None when it matches the
device's most recent row: that boot is already recorded.
Only the LATEST row is compared, never "any row with this boot_count": the
firmware's lifetime counter gets reset (reflash / telemetry reset), so the
same boot_count legitimately appears again for a later, different boot. A
genuine new boot always differs from the latest row — the counter either
moved forward or was reset to a lower number."""
same boot_count legitimately appears again for a later, different boot."""
async with AsyncSessionLocal() as session:
if boot_count is not None:
if skip_if_latest and boot_count is not None:
latest = await session.execute(
text("""
SELECT boot_count, reset_reason FROM device_boot_events