fix(mqtt): treat broker-replayed retained messages as stale state
The firmware publishes status/heartbeat, system/alerts, system/info and
status/playback with retain=true. On every backend (re)connect - every
restart and every uvicorn --reload - the broker replays the last message on
each of those topics for every device that ever connected. We handled those
replays as if they had just happened:
- heartbeats: a row with received_at=now() for every device, so devices
that have been dead for months showed ONLINE for 90s after each restart
and got pinged. ~770k such rows exist locally.
- boot_report: the last boot logged again as a new reboot (the phantom
PANIC entries on the Health tab).
- alerts / other info events: logged again as new occurrences.
MQTT delivers retain=1 only for replays caused by a new subscription; live
publishes always arrive with retain=0. The flag is now passed through to the
handlers and the WS broadcast:
- heartbeat: replays are not stored. A live heartbeat is.
- {"state":"offline"} heartbeat (LWT / graceful disconnect) is no longer
stored as a sign of life. It marks the device offline immediately in
a small in-memory set (mqtt/presence.py) used by /mqtt/status and the ping
loop; a later live heartbeat clears it. Replayed offline markers also mark
offline, since a retained message is the device's last word.
- boot_report: live -> always a new boot. Replay -> stored only if it
differs from the device's latest boot row (i.e. we missed it while down).
- alerts: replay still syncs the current-alert row; history gets a row on a
live alert (even an identical repeat - faults recur) or on a replay that
changes state. Replaces the 98dd16b rule that dropped identical live alerts.
- other info events: replays are not logged.
- Frontend (DeviceList, DeviceDetail, LogsTab) ignores retained WS messages
for live updates, and flips a device offline on the offline marker instead
of marking it online.
Verified locally after a backend restart: only the 7 actually-live devices
got new heartbeat rows (none from the replays), and no boot/alert/info rows
were created.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
@@ -444,20 +444,17 @@ async def insert_boot_event(device_serial: str, boot_count: int | None,
|
||||
crash_task: str | None = None,
|
||||
crash_pc: int | None = None,
|
||||
crash_exc_cause: int | None = None,
|
||||
crash_exc_vaddr: int | None = None) -> int | None:
|
||||
"""Insert a boot event, unless it is a repeat of the device's most recent
|
||||
one (returns None then). boot_report is published retained on system/info,
|
||||
so the broker redelivers the last one every time the backend (re)subscribes
|
||||
— without this check each backend restart logged the device's last boot
|
||||
(crash detail included) again as a brand-new event.
|
||||
crash_exc_vaddr: int | None = None,
|
||||
skip_if_latest: bool = False) -> int | None:
|
||||
"""Insert a boot event. With skip_if_latest=True (a retained replay of the
|
||||
device's last boot_report), skip it and return None when it matches the
|
||||
device's most recent row: that boot is already recorded.
|
||||
|
||||
Only the LATEST row is compared, never "any row with this boot_count": the
|
||||
firmware's lifetime counter gets reset (reflash / telemetry reset), so the
|
||||
same boot_count legitimately appears again for a later, different boot. A
|
||||
genuine new boot always differs from the latest row — the counter either
|
||||
moved forward or was reset to a lower number."""
|
||||
same boot_count legitimately appears again for a later, different boot."""
|
||||
async with AsyncSessionLocal() as session:
|
||||
if boot_count is not None:
|
||||
if skip_if_latest and boot_count is not None:
|
||||
latest = await session.execute(
|
||||
text("""
|
||||
SELECT boot_count, reset_reason FROM device_boot_events
|
||||
|
||||
Reference in New Issue
Block a user