feat(mqtt): store F-070 crash detail and group crashes fleet-wide
Firmware F-070 adds a fuller `crash` object (backtrace, backtrace_corrupted, elf_sha256) and a new `pre_crash` snapshot (uptime, heap stats, optional abort_msg) to boot_report and to telemetry.get_boot_history entries. - Migration c9d0e1f2a3b4: device_boot_events gains crash / pre_crash JSONB, stored verbatim so later additive fields need no migration. The legacy crash_* columns are still filled; old rows and old firmware are unchanged. JSON is bound as CAST(CAST(:x AS TEXT) AS JSONB): a bare JSONB cast makes asyncpg JSON-encode the already-encoded string a second time. - boot_report ingestion passes both objects through. - telemetry.get_boot_history replies (control/ack) are merged into boot history whoever sent the command: an entry matches an existing row on boot_count + reset_reason + time within 30 min (boot_count alone is not unique - the counter gets reset), and only fills crash/pre_crash the row lacks; unmatched entries are boots we never saw live and are inserted at the device's timestamp; entries without ts are skipped. - GET /api/mqtt/crash-groups: fault boots grouped by abort_msg with hex addresses stripped, else task + exception cause, else reset reason. When abort_msg is present, pc/exc_cause describe abort() itself and are ignored for grouping. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This commit is contained in:
+16
-5
@@ -163,17 +163,19 @@ async def _handle_boot_report(serial: str, payload: dict, retained: bool = False
|
||||
info event as { "type": ..., "payload": {...} }.
|
||||
"""
|
||||
data = payload.get("payload", {})
|
||||
crash = data.get("crash") or {}
|
||||
# F-070: `crash` (coredump summary, now incl. backtrace/elf_sha256) and
|
||||
# `pre_crash` (heap/uptime snapshot + optional abort_msg) are stored
|
||||
# verbatim. Older firmware omits them / sends the 4-field crash only.
|
||||
crash = data.get("crash") if isinstance(data.get("crash"), dict) else None
|
||||
pre_crash = data.get("pre_crash") if isinstance(data.get("pre_crash"), dict) else None
|
||||
await db.insert_boot_event(
|
||||
device_serial=serial,
|
||||
boot_count=data.get("boot_count"),
|
||||
reset_reason=data.get("reset_reason"),
|
||||
is_fault=bool(data.get("is_fault", False)),
|
||||
free_heap=data.get("free_heap"),
|
||||
crash_task=crash.get("task"),
|
||||
crash_pc=crash.get("pc"),
|
||||
crash_exc_cause=crash.get("exc_cause"),
|
||||
crash_exc_vaddr=crash.get("exc_vaddr"),
|
||||
crash=crash,
|
||||
pre_crash=pre_crash,
|
||||
# A live boot_report is always a new boot. A retained replay is the
|
||||
# device's last boot, recorded only if we missed it while down.
|
||||
skip_if_latest=retained,
|
||||
@@ -252,6 +254,15 @@ async def _handle_ack(serial: str, payload: dict):
|
||||
await db.insert_ping_sample(device_serial=serial, rtt_ms=rtt_ms)
|
||||
return
|
||||
|
||||
# The device's own SD boot log — merged into device_boot_events whoever
|
||||
# asked for it (Health tab "Sync from device", API reference, Control tab),
|
||||
# so crash detail for boots we never saw live isn't lost.
|
||||
if payload.get("type") == "telemetry.get_boot_history" and status == "SUCCESS":
|
||||
boots = (payload.get("data") or {}).get("boots")
|
||||
if isinstance(boots, list):
|
||||
result = await db.merge_device_boot_history(serial, boots)
|
||||
logger.info(f"Merged device boot history for {serial}: {result}")
|
||||
|
||||
pending = await db.get_pending_command(serial)
|
||||
if pending:
|
||||
cmd_status = "success" if status == "SUCCESS" else "error"
|
||||
|
||||
Reference in New Issue
Block a user