fix(mqtt): treat broker-replayed retained messages as stale state
The firmware publishes status/heartbeat, system/alerts, system/info and
status/playback with retain=true. On every backend (re)connect - every
restart and every uvicorn --reload - the broker replays the last message on
each of those topics for every device that ever connected. We handled those
replays as if they had just happened:
- heartbeats: a row with received_at=now() for every device, so devices
that have been dead for months showed ONLINE for 90s after each restart
and got pinged. ~770k such rows exist locally.
- boot_report: the last boot logged again as a new reboot (the phantom
PANIC entries on the Health tab).
- alerts / other info events: logged again as new occurrences.
MQTT delivers retain=1 only for replays caused by a new subscription; live
publishes always arrive with retain=0. The flag is now passed through to the
handlers and the WS broadcast:
- heartbeat: replays are not stored. A live heartbeat is.
- {"state":"offline"} heartbeat (LWT / graceful disconnect) is no longer
stored as a sign of life. It marks the device offline immediately in
a small in-memory set (mqtt/presence.py) used by /mqtt/status and the ping
loop; a later live heartbeat clears it. Replayed offline markers also mark
offline, since a retained message is the device's last word.
- boot_report: live -> always a new boot. Replay -> stored only if it
differs from the device's latest boot row (i.e. we missed it while down).
- alerts: replay still syncs the current-alert row; history gets a row on a
live alert (even an identical repeat - faults recur) or on a replay that
changes state. Replaces the 98dd16b rule that dropped identical live alerts.
- other info events: replays are not logged.
- Frontend (DeviceList, DeviceDetail, LogsTab) ignores retained WS messages
for live updates, and flips a device offline on the offline marker instead
of marking it online.
Verified locally after a backend restart: only the 7 actually-live devices
got new heartbeat rows (none from the replays), and no boot/alert/info rows
were created.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
+13
-4
@@ -99,10 +99,17 @@ class MqttManager:
|
||||
|
||||
serial = parts[1]
|
||||
topic_type = "/".join(parts[2:])
|
||||
# The broker sets retain=1 only on a message it replays from its
|
||||
# retained store because we just (re)subscribed: a stale snapshot
|
||||
# of the device's last state, not something that just happened.
|
||||
# Live publishes always arrive with retain=0, even on topics the
|
||||
# firmware publishes retained. Every backend restart (and every
|
||||
# uvicorn --reload) replays these for every device.
|
||||
retained = bool(msg.retain)
|
||||
|
||||
if self._loop and self._loop.is_running():
|
||||
asyncio.run_coroutine_threadsafe(
|
||||
self._process_message(serial, topic_type, payload, topic),
|
||||
self._process_message(serial, topic_type, payload, topic, retained),
|
||||
self._loop,
|
||||
)
|
||||
except json.JSONDecodeError:
|
||||
@@ -111,15 +118,16 @@ class MqttManager:
|
||||
logger.error(f"Error processing MQTT message: {e}")
|
||||
|
||||
async def _process_message(self, serial: str, topic_type: str,
|
||||
payload: dict, raw_topic: str):
|
||||
payload: dict, raw_topic: str, retained: bool = False):
|
||||
from mqtt.logger import handle_message
|
||||
await handle_message(serial, topic_type, payload)
|
||||
await handle_message(serial, topic_type, payload, retained=retained)
|
||||
|
||||
ws_data = {
|
||||
"type": topic_type,
|
||||
"device_serial": serial,
|
||||
"payload": payload,
|
||||
"topic": raw_topic,
|
||||
"retained": retained,
|
||||
}
|
||||
await self._broadcast_ws(ws_data)
|
||||
|
||||
@@ -170,6 +178,7 @@ class MqttManager:
|
||||
|
||||
async def _ping_online_devices(self):
|
||||
import database as db
|
||||
from mqtt import presence
|
||||
heartbeats = await db.get_latest_heartbeats()
|
||||
now = time.time()
|
||||
for hb in heartbeats:
|
||||
@@ -179,7 +188,7 @@ class MqttManager:
|
||||
age = now - received.timestamp()
|
||||
except (ValueError, TypeError, KeyError):
|
||||
continue
|
||||
if age > PING_ONLINE_WINDOW_SECONDS:
|
||||
if age > PING_ONLINE_WINDOW_SECONDS or presence.is_marked_offline(hb["device_serial"]):
|
||||
continue
|
||||
self.publish_command(
|
||||
device_serial=hb["device_serial"],
|
||||
|
||||
Reference in New Issue
Block a user