The firmware publishes status/heartbeat, system/alerts, system/info and
status/playback with retain=true. On every backend (re)connect - every
restart and every uvicorn --reload - the broker replays the last message on
each of those topics for every device that ever connected. We handled those
replays as if they had just happened:
- heartbeats: a row with received_at=now() for every device, so devices
that have been dead for months showed ONLINE for 90s after each restart
and got pinged. ~770k such rows exist locally.
- boot_report: the last boot logged again as a new reboot (the phantom
PANIC entries on the Health tab).
- alerts / other info events: logged again as new occurrences.
MQTT delivers retain=1 only for replays caused by a new subscription; live
publishes always arrive with retain=0. The flag is now passed through to the
handlers and the WS broadcast:
- heartbeat: replays are not stored. A live heartbeat is.
- {"state":"offline"} heartbeat (LWT / graceful disconnect) is no longer
stored as a sign of life. It marks the device offline immediately in
a small in-memory set (mqtt/presence.py) used by /mqtt/status and the ping
loop; a later live heartbeat clears it. Replayed offline markers also mark
offline, since a retained message is the device's last word.
- boot_report: live -> always a new boot. Replay -> stored only if it
differs from the device's latest boot row (i.e. we missed it while down).
- alerts: replay still syncs the current-alert row; history gets a row on a
live alert (even an identical repeat - faults recur) or on a replay that
changes state. Replaces the 98dd16b rule that dropped identical live alerts.
- other info events: replays are not logged.
- Frontend (DeviceList, DeviceDetail, LogsTab) ignores retained WS messages
for live updates, and flips a device offline on the offline marker instead
of marking it online.
Verified locally after a backend restart: only the 7 actually-live devices
got new heartbeat rows (none from the replays), and no boot/alert/info rows
were created.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
27 lines
878 B
Python
27 lines
878 B
Python
"""Devices that announced they went offline.
|
|
|
|
The firmware's LWT (unclean disconnect) and its graceful-disconnect message both
|
|
publish {"state": "offline", "ok": false} to status/heartbeat, retained. Online
|
|
status is otherwise "a heartbeat newer than 90s", which would keep showing a
|
|
device that just dropped as online for up to 90s. This remembers the marker so
|
|
/mqtt/status and the ping loop treat the device as offline right away.
|
|
|
|
In-memory only, which is enough: on backend start the broker replays each
|
|
device's retained heartbeat, so a device whose last word was "offline" gets
|
|
re-marked, and any live heartbeat clears the mark.
|
|
"""
|
|
|
|
_offline: set[str] = set()
|
|
|
|
|
|
def mark_offline(serial: str) -> None:
|
|
_offline.add(serial)
|
|
|
|
|
|
def mark_alive(serial: str) -> None:
|
|
_offline.discard(serial)
|
|
|
|
|
|
def is_marked_offline(serial: str) -> bool:
|
|
return serial in _offline
|