feat(mqtt): store F-070 crash detail and group crashes fleet-wide

Firmware F-070 adds a fuller `crash` object (backtrace, backtrace_corrupted,
elf_sha256) and a new `pre_crash` snapshot (uptime, heap stats, optional
abort_msg) to boot_report and to telemetry.get_boot_history entries.

- Migration c9d0e1f2a3b4: device_boot_events gains crash / pre_crash JSONB,
  stored verbatim so later additive fields need no migration. The legacy
  crash_* columns are still filled; old rows and old firmware are unchanged.
  JSON is bound as CAST(CAST(:x AS TEXT) AS JSONB): a bare JSONB cast makes
  asyncpg JSON-encode the already-encoded string a second time.
- boot_report ingestion passes both objects through.
- telemetry.get_boot_history replies (control/ack) are merged into boot
  history whoever sent the command: an entry matches an existing row on
  boot_count + reset_reason + time within 30 min (boot_count alone is not
  unique - the counter gets reset), and only fills crash/pre_crash the row
  lacks; unmatched entries are boots we never saw live and are inserted at
  the device's timestamp; entries without ts are skipped.
- GET /api/mqtt/crash-groups: fault boots grouped by abort_msg with hex
  addresses stripped, else task + exception cause, else reset reason. When
  abort_msg is present, pc/exc_cause describe abort() itself and are
  ignored for grouping.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
2026-09-30 16:28:40 +03:00
co-authored by Claude Opus 5.5
parent 46c5c0a846
commit 963516ece0
7 changed files with 365 additions and 25 deletions
+16 -5
View File
@@ -163,17 +163,19 @@ async def _handle_boot_report(serial: str, payload: dict, retained: bool = False
info event as { "type": ..., "payload": {...} }.
"""
data = payload.get("payload", {})
crash = data.get("crash") or {}
# F-070: `crash` (coredump summary, now incl. backtrace/elf_sha256) and
# `pre_crash` (heap/uptime snapshot + optional abort_msg) are stored
# verbatim. Older firmware omits them / sends the 4-field crash only.
crash = data.get("crash") if isinstance(data.get("crash"), dict) else None
pre_crash = data.get("pre_crash") if isinstance(data.get("pre_crash"), dict) else None
await db.insert_boot_event(
device_serial=serial,
boot_count=data.get("boot_count"),
reset_reason=data.get("reset_reason"),
is_fault=bool(data.get("is_fault", False)),
free_heap=data.get("free_heap"),
crash_task=crash.get("task"),
crash_pc=crash.get("pc"),
crash_exc_cause=crash.get("exc_cause"),
crash_exc_vaddr=crash.get("exc_vaddr"),
crash=crash,
pre_crash=pre_crash,
# A live boot_report is always a new boot. A retained replay is the
# device's last boot, recorded only if we missed it while down.
skip_if_latest=retained,
@@ -252,6 +254,15 @@ async def _handle_ack(serial: str, payload: dict):
await db.insert_ping_sample(device_serial=serial, rtt_ms=rtt_ms)
return
# The device's own SD boot log — merged into device_boot_events whoever
# asked for it (Health tab "Sync from device", API reference, Control tab),
# so crash detail for boots we never saw live isn't lost.
if payload.get("type") == "telemetry.get_boot_history" and status == "SUCCESS":
boots = (payload.get("data") or {}).get("boots")
if isinstance(boots, list):
result = await db.merge_device_boot_history(serial, boots)
logger.info(f"Merged device boot history for {serial}: {result}")
pending = await db.get_pending_command(serial)
if pending:
cmd_status = "success" if status == "SUCCESS" else "error"