fix(mqtt): treat broker-replayed retained messages as stale state

The firmware publishes status/heartbeat, system/alerts, system/info and
status/playback with retain=true. On every backend (re)connect - every
restart and every uvicorn --reload - the broker replays the last message on
each of those topics for every device that ever connected. We handled those
replays as if they had just happened:

- heartbeats: a row with received_at=now() for every device, so devices
  that have been dead for months showed ONLINE for 90s after each restart
  and got pinged. ~770k such rows exist locally.
- boot_report: the last boot logged again as a new reboot (the phantom
  PANIC entries on the Health tab).
- alerts / other info events: logged again as new occurrences.

MQTT delivers retain=1 only for replays caused by a new subscription; live
publishes always arrive with retain=0. The flag is now passed through to the
handlers and the WS broadcast:

- heartbeat: replays are not stored. A live heartbeat is.
- {"state":"offline"} heartbeat (LWT / graceful disconnect) is no longer
  stored as a sign of life. It marks the device offline immediately in
  a small in-memory set (mqtt/presence.py) used by /mqtt/status and the ping
  loop; a later live heartbeat clears it. Replayed offline markers also mark
  offline, since a retained message is the device's last word.
- boot_report: live -> always a new boot. Replay -> stored only if it
  differs from the device's latest boot row (i.e. we missed it while down).
- alerts: replay still syncs the current-alert row; history gets a row on a
  live alert (even an identical repeat - faults recur) or on a replay that
  changes state. Replaces the 98dd16b rule that dropped identical live alerts.
- other info events: replays are not logged.
- Frontend (DeviceList, DeviceDetail, LogsTab) ignores retained WS messages
  for live updates, and flips a device offline on the offline marker instead
  of marking it online.

Verified locally after a backend restart: only the 7 actually-live devices
got new heartbeat rows (none from the replays), and no boot/alert/info rows
were created.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-30 16:01:41 +03:00
co-authored by Claude Opus 5.5
parent 18e57c6e5f
commit 46c5c0a846
8 changed files with 111 additions and 27 deletions
+2 -1
View File
@@ -11,6 +11,7 @@ from mqtt.models import (
LatestDiagnosticsEntry, LatestPingEntry, DeviceReportListResponse,
)
from mqtt.client import mqtt_manager
from mqtt import presence
import database as db
from datetime import datetime, timezone
@@ -40,7 +41,7 @@ async def get_all_device_status(
devices.append(DeviceMqttStatus(
device_serial=hb["device_serial"],
online=seconds_ago < 90,
online=seconds_ago < 90 and not presence.is_marked_offline(hb["device_serial"]),
last_heartbeat=HeartbeatEntry(**hb),
seconds_since_heartbeat=seconds_ago,
last_alert_event=AlertEventEntry(**alert_event) if alert_event else None,