fix(mqtt): treat broker-replayed retained messages as stale state

The firmware publishes status/heartbeat, system/alerts, system/info and
status/playback with retain=true. On every backend (re)connect - every
restart and every uvicorn --reload - the broker replays the last message on
each of those topics for every device that ever connected. We handled those
replays as if they had just happened:

- heartbeats: a row with received_at=now() for every device, so devices
  that have been dead for months showed ONLINE for 90s after each restart
  and got pinged. ~770k such rows exist locally.
- boot_report: the last boot logged again as a new reboot (the phantom
  PANIC entries on the Health tab).
- alerts / other info events: logged again as new occurrences.

MQTT delivers retain=1 only for replays caused by a new subscription; live
publishes always arrive with retain=0. The flag is now passed through to the
handlers and the WS broadcast:

- heartbeat: replays are not stored. A live heartbeat is.
- {"state":"offline"} heartbeat (LWT / graceful disconnect) is no longer
  stored as a sign of life. It marks the device offline immediately in
  a small in-memory set (mqtt/presence.py) used by /mqtt/status and the ping
  loop; a later live heartbeat clears it. Replayed offline markers also mark
  offline, since a retained message is the device's last word.
- boot_report: live -> always a new boot. Replay -> stored only if it
  differs from the device's latest boot row (i.e. we missed it while down).
- alerts: replay still syncs the current-alert row; history gets a row on a
  live alert (even an identical repeat - faults recur) or on a replay that
  changes state. Replaces the 98dd16b rule that dropped identical live alerts.
- other info events: replays are not logged.
- Frontend (DeviceList, DeviceDetail, LogsTab) ignores retained WS messages
  for live updates, and flips a device offline on the offline marker instead
  of marking it online.

Verified locally after a backend restart: only the 7 actually-live devices
got new heartbeat rows (none from the replays), and no boot/alert/info rows
were created.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-30 16:01:41 +03:00
co-authored by Claude Opus 5.5
parent 18e57c6e5f
commit 46c5c0a846
8 changed files with 111 additions and 27 deletions
+13 -4
View File
@@ -99,10 +99,17 @@ class MqttManager:
serial = parts[1]
topic_type = "/".join(parts[2:])
# The broker sets retain=1 only on a message it replays from its
# retained store because we just (re)subscribed: a stale snapshot
# of the device's last state, not something that just happened.
# Live publishes always arrive with retain=0, even on topics the
# firmware publishes retained. Every backend restart (and every
# uvicorn --reload) replays these for every device.
retained = bool(msg.retain)
if self._loop and self._loop.is_running():
asyncio.run_coroutine_threadsafe(
self._process_message(serial, topic_type, payload, topic),
self._process_message(serial, topic_type, payload, topic, retained),
self._loop,
)
except json.JSONDecodeError:
@@ -111,15 +118,16 @@ class MqttManager:
logger.error(f"Error processing MQTT message: {e}")
async def _process_message(self, serial: str, topic_type: str,
payload: dict, raw_topic: str):
payload: dict, raw_topic: str, retained: bool = False):
from mqtt.logger import handle_message
await handle_message(serial, topic_type, payload)
await handle_message(serial, topic_type, payload, retained=retained)
ws_data = {
"type": topic_type,
"device_serial": serial,
"payload": payload,
"topic": raw_topic,
"retained": retained,
}
await self._broadcast_ws(ws_data)
@@ -170,6 +178,7 @@ class MqttManager:
async def _ping_online_devices(self):
import database as db
from mqtt import presence
heartbeats = await db.get_latest_heartbeats()
now = time.time()
for hb in heartbeats:
@@ -179,7 +188,7 @@ class MqttManager:
age = now - received.timestamp()
except (ValueError, TypeError, KeyError):
continue
if age > PING_ONLINE_WINDOW_SECONDS:
if age > PING_ONLINE_WINDOW_SECONDS or presence.is_marked_offline(hb["device_serial"]):
continue
self.publish_command(
device_serial=hb["device_serial"],