1. Topic

Log levels, rate limiting, event ring buffers, breadcrumbs, fault snapshots, binary events, UART/MQTT/CAN telemetry and the diag/events/health CLI. Main idea: logging can itself cause latency, buffer overflow and watchdog resets.

2. Why this matters in a project

You need to know which AT command preceded an EC25 failure, why MQTT reconnected, whether queues drop events, whether traffic-light inputs are noisy, whether I2C timed out, UART overran or CAN entered bus-off, and which task failed to make progress before reset.

3. Theory

Diagnostic layers:

text
1. Counters - fast and inexpensive.
2. Event ring buffer - recent events.
3. Text logs - for humans.
4. Fault snapshot - before reset/fatal fault.
5. Core dump - post-mortem tasks/stacks/registers.

Levels:

text
ERROR - unrecovered error, data loss, reset.
WARN - recovered fault/degraded mode.
INFO - boot, connected, OTA start, mode change.
DEBUG - state machine internals.
VERBOSE - byte dumps/raw samples.

The fast path does not print text. It updates counters and an event ring. Rate limiting:

c
typedef struct {
    int64_t last_us;
    uint32_t suppressed;
} log_rl_t;
static bool log_rate_allow(log_rl_t *rl, int64_t now_us, int64_t period_us)
{
    if ((now_us - rl->last_us) >= period_us) {
        rl->last_us = now_us;
        return true;
    }
    rl->suppressed++;
    return false;
}

Event ring:

c
typedef struct {
    int64_t ts_us;
    uint16_t module;
    uint16_t event;
    uint32_t a, b, c;
} diag_event_t;

Do not send network logs directly from the failure site. A module records an event/counter; logger_task sends on a best-effort basis, dropping DEBUG/INFO first under overload. The diag CLI should show heap, stack watermarks, queue fill/drops, DMA overruns, reconnect count and last reset reason.

4. Common mistakes

  1. Logging in the fast path.
  2. Logging an error without a counter and context.
  3. Omitting rate limits.
  4. Sending logs over the failing channel.
  5. Assuming DEBUG is harmless.
  6. Dumping large buffers without limits.
  7. Failing to save breadcrumbs before a watchdog reset.

5. Practical task

Create LOGGING_DIAGNOSTICS_POLICY.md:

markdown
# Logging and diagnostics policy
Main rule:
Logs are not part of fast path timing.
Rules:
1. Fast path updates counters/events, not text logs.
2. Every repeated error has counter and rate limit.
3. ERROR/WARN logs include context.
4. DEBUG/VERBOSE are disabled in production by default.
5. Hex dumps are length-limited and debug-only.
6. Network log transport must not block critical tasks.
7. Last diagnostic events are stored in ring buffer.
8. Fault snapshot is saved before intentional reset.
9. CLI exposes counters, queues, health, events and reset reason.

Enums:

c
typedef enum {
    DIAG_MOD_SYSTEM = 1,
    DIAG_MOD_INPUT,
    DIAG_MOD_PHASE,
    DIAG_MOD_MODEM,
    DIAG_MOD_MQTT,
    DIAG_MOD_I2C,
    DIAG_MOD_UART,
    DIAG_MOD_CAN,
    DIAG_MOD_LOGGER,
} diag_module_t;

6. Further reading and experiments

ESP-IDF Logging Library, Core Dump/Fatal Errors and Heap Debugging; ST AN4989 on SWV/ITM/SWO.

An error occurs in the input fast path. What should happen first?

Exercise

UART runs at 115200 baud with 8N1. A logger generates 400 lines/s of 40 bytes each, including newline. Ignoring other overhead, is the link sufficient? Choose a bounded overload policy.

Self-check criteria: Show 10 bits/byte, 11520 and 16000 bytes/s; propose a bounded queue and dropped-log accounting.

Show the supplied answer

8N1 needs 10 bits/byte: 115200/10=11520 bytes/s. Logging requires 400×40=16000 bytes/s, so it is insufficient: an unbounded queue would grow. Use a bounded queue, rate limiting, aggregate counters and DEBUG/INFO drop accounting. Preserve important fault events under a separate bounded policy.