Commit Graph
7 Commits
Author SHA1 Message Date
bonaminandClaude Opus 5.5 963516ece0 feat(mqtt): store F-070 crash detail and group crashes fleet-wide
Firmware F-070 adds a fuller `crash` object (backtrace, backtrace_corrupted,
elf_sha256) and a new `pre_crash` snapshot (uptime, heap stats, optional
abort_msg) to boot_report and to telemetry.get_boot_history entries.

- Migration c9d0e1f2a3b4: device_boot_events gains crash / pre_crash JSONB,
  stored verbatim so later additive fields need no migration. The legacy
  crash_* columns are still filled; old rows and old firmware are unchanged.
  JSON is bound as CAST(CAST(:x AS TEXT) AS JSONB): a bare JSONB cast makes
  asyncpg JSON-encode the already-encoded string a second time.
- boot_report ingestion passes both objects through.
- telemetry.get_boot_history replies (control/ack) are merged into boot
  history whoever sent the command: an entry matches an existing row on
  boot_count + reset_reason + time within 30 min (boot_count alone is not
  unique - the counter gets reset), and only fills crash/pre_crash the row
  lacks; unmatched entries are boots we never saw live and are inserted at
  the device's timestamp; entries without ts are skipped.
- GET /api/mqtt/crash-groups: fault boots grouped by abort_msg with hex
  addresses stripped, else task + exception cause, else reset reason. When
  abort_msg is present, pc/exc_cause describe abort() itself and are
  ignored for grouping.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 16:28:40 +03:00
bonaminandClaude Opus 5.5 46c5c0a846 fix(mqtt): treat broker-replayed retained messages as stale state
The firmware publishes status/heartbeat, system/alerts, system/info and
status/playback with retain=true. On every backend (re)connect - every
restart and every uvicorn --reload - the broker replays the last message on
each of those topics for every device that ever connected. We handled those
replays as if they had just happened:

- heartbeats: a row with received_at=now() for every device, so devices
  that have been dead for months showed ONLINE for 90s after each restart
  and got pinged. ~770k such rows exist locally.
- boot_report: the last boot logged again as a new reboot (the phantom
  PANIC entries on the Health tab).
- alerts / other info events: logged again as new occurrences.

MQTT delivers retain=1 only for replays caused by a new subscription; live
publishes always arrive with retain=0. The flag is now passed through to the
handlers and the WS broadcast:

- heartbeat: replays are not stored. A live heartbeat is.
- {"state":"offline"} heartbeat (LWT / graceful disconnect) is no longer
  stored as a sign of life. It marks the device offline immediately in
  a small in-memory set (mqtt/presence.py) used by /mqtt/status and the ping
  loop; a later live heartbeat clears it. Replayed offline markers also mark
  offline, since a retained message is the device's last word.
- boot_report: live -> always a new boot. Replay -> stored only if it
  differs from the device's latest boot row (i.e. we missed it while down).
- alerts: replay still syncs the current-alert row; history gets a row on a
  live alert (even an identical repeat - faults recur) or on a replay that
  changes state. Replaces the 98dd16b rule that dropped identical live alerts.
- other info events: replays are not logged.
- Frontend (DeviceList, DeviceDetail, LogsTab) ignores retained WS messages
  for live updates, and flips a device offline on the offline marker instead
  of marking it online.

Verified locally after a backend restart: only the 7 actually-live devices
got new heartbeat rows (none from the replays), and no boot/alert/info rows
were created.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 16:01:41 +03:00
bonaminandClaude Opus 5.5 18e57c6e5f fix(mqtt): dedupe boot_report against the latest row, not any boot_count
98dd16b skipped a boot_report if ANY earlier row had the same boot_count.
That is wrong: the firmware's lifetime boot counter gets reset (reflash /
telemetry reset), and the data shows counts 1-4 recurring in July and again
in September. With that rule a real later boot reusing a number would be
dropped forever.

A retained redelivery is always a copy of the device's most recent boot, so
compare only against the latest row (boot_count + reset_reason). A genuine
new boot always differs from it - the counter moves forward or was reset.

Note: the one-off cleanup run on 2026-09-30 used the same wrong
(serial, boot_count) key and deleted some genuine boot rows along with the
redelivery duplicates; see the session notes / heartbeat-based reboot list.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:55:26 +03:00
bonaminandClaude Opus 5.5 98dd16b597 fix(mqtt): stop logging retained boot_report/alerts as new events
The firmware publishes boot_report on system/info and alerts on
system/alerts with retain=true. Every time the backend (re)connects - which
under uvicorn --reload means every backend file save - the broker redelivers
the last retained message and we inserted it again with occurred_at=now().
Result: the Health tab showed fresh PANIC boots and "Device reset due to
fault" alerts for a device that had been up for 4 days.

- insert_boot_event skips the insert when a row with the same
  (device_serial, boot_count) already exists. boot_count is the firmware's
  lifetime counter, so it uniquely identifies a boot.
- upsert_alert only writes when state/message actually changed and returns
  whether it did; the alert-event history row is only added on a change.
  A redelivered identical alert no longer bumps updated_at either.

Existing duplicate rows are not touched by this commit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 15:48:48 +03:00
bonaminandClaude Sonnet 5 7c533b9245 feat(mqtt): add device health telemetry and migrate to v2 topic spec
Two efforts that landed together because the v2 topic work extends
tables the health-telemetry effort added days earlier in the same
files/functions, making them impractical to separate cleanly:

Health/diagnostics telemetry (schema, Jul 13-17):
- New Postgres tables: device_alert_events, device_boot_events,
  device_ping_samples, device_diagnostics_reports, plus a `source`
  column on device_logs to distinguish log origins
- Query/service layer in pg_mqtt.py and database/__init__.py for
  inserting and listing this history, plus a "latest metrics" endpoint
  combining most-recent diagnostics + ping RTT per device
- mqtt/router.py gains list endpoints for alert/boot/ping/diagnostics
  history, consumed by the upcoming Health tab

MQTT v2 topic migration (Sep 21):
- Heartbeat payload flattened per vesper_mqtt_topic_spec_v2.md, adding
  rssi/free_heap/state/ok fields
- Command replies move to control/ack, device-initiated events to
  control/reports; mqtt/client.py subscribes to the new topic set and
  runs a ping_loop (wired up in main.py) for RTT sampling
- mqtt/logger.py and pg_mqtt.py updated to parse and persist the new
  payload shape alongside the legacy fields for backwards compatibility

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-21 18:24:36 +03:00
bonamin 2ef199e4c5 fix: login issue 2026-04-17 16:01:50 +03:00
bonamin a605143c5d Phase 5 of Migration 2026-04-17 15:51:27 +03:00