1. Today's topic

We move from “it is fast on average” to the question:

text
what is the maximum time from a hardware event to the response?

Topics:

c
deadline;
period;
latency;
jitter;
WCET;
ISR latency;
scheduler latency;
priority inversion;
response-time analysis;
DMA buffer deadlines;
Flash/NVS interference;
latency histogram;
deadline misses.

The main idea: average CPU load does not prove that deadlines are met. At 20% CPU load, one long critical section or Flash write can still cause a missed deadline.

2. Why this matters

Phase inputs

text
signal -> opto -> GPIO/ADC -> ISR/DMA -> queue -> input_task -> filter -> phase FSM

Total delay:

text
Tdecision = Thardware + Tinterrupt + Tqueue + Tscheduler + Tprocessing + Tfilter

Separate intentional debounce time from unwanted RTOS delay.

EC25 UART

115200 8N1:

text
Tbyte ~= 10 / 115200 = 86.8 us
256 bytes ~= 22.2 ms

If the UART consumer does not run for longer than it takes the ring buffer to fill, data is lost.

ADC DMA

text
Tblock = Nsamples / Faggregate

For a half-buffer of 128 samples and an aggregate rate of 32 ksample/s:

text
Tblock = 4 ms

Processing must finish faster, preferably with a margin.

3. Theory

Latency components

text
Ltotal = Lirq_disabled + Cisr + Lwake + Lschedule + Ctask + Bresource + Cout

Each component needs a budget, a measured maximum and a deadline-miss counter.

Response-time analysis

For fixed priorities on one core:

text
R_i(n+1) = C_i + B_i + J_i + sum(ceil((R_i(n)+J_j)/T_j) * C_j)

If R > D, the task misses its deadline.

Priorities

Priority is determined by the deadline, not by business importance.

text
The UART buffer will overflow in 10 ms -> UART service above MQTT.
The ADC half-buffer will be overwritten in 4 ms -> ADC service above logger.

A high-priority task must spend most of its time BLOCKED, rather than spinning in a polling loop.

ESP32 SMP

On a dual-core ESP32, account for:

text
affinity;
system tasks;
spinlocks;
heap locks;
inter-core queues;
Flash operations;
shared memory bus.

Pinning helps only if shared resources do not become a bottleneck.

FreeRTOS tick versus hardware time

vTaskDelayUntil() is better for periodic tasks than vTaskDelay(). For microsecond timestamps, use esp_timer_get_time() or a hardware timer/input capture.

Flash/NVS

Flash erase/write can delay task execution. Do not assume that a low-priority storage task cannot affect a high-priority path.

Deadline miss

A deadline miss is a system event, not merely a warning:

c
phase data -> DEGRADED;
UART loss -> resync parser;
ADC block miss -> drop block;
CAN FIFO overrun -> fault counter;
MQTT delay -> metric only.

4. Common mistakes

  • Evaluating only average CPU load.
  • Giving every important task the same high priority.
  • Placing the logger above UART RX.
  • Running a parser in an ISR.
  • Using vTaskDelay() as a precise period.
  • Ignoring the DMA block period.
  • Holding a mutex during I/O.
  • Assuming two cores automatically improve real-time performance.
  • Leaving a time-critical task unpinned on SMP without analysis.
  • Excluding Flash erase/write from the threat model.
  • Taking the timestamp only in the task.
  • Storing only an average, without a histogram.
  • Not counting ISR drops.
  • Printing every deadline miss synchronously.
  • Treating a watchdog as a micro-deadline monitor.

5. A practical task for 30–60 minutes

Create LATENCY_BUDGET.md:

markdown
# Latency budget
| Pipeline | Period/min interval | Deadline | WCET target |
|---|---:|---:|---:|
| Phase GPIO ISR -> task | 2 ms | 500 us | 100 us |
| Phase raw -> FSM | 2 ms | 1 ms | 300 us |
| EC25 UART service | 5 ms | 2 ms | 500 us |
| ADC half-buffer | 4 ms | 4 ms | 1 ms |
| MQTT command | - | 500 ms | 20 ms |
| Logger | - | 2 s | 100 ms |

Add instrumentation:

c
ISR timestamp;
task-start timestamp;
processing-finished timestamp;
queue high water;
drops;
deadline misses;
histogram.

CLI:

text
latency phase
events=10042
irq_drops=0
scheduler:
min=8us
avg=24us
max=184us
deadline=500us
misses=0
end_to_end:
min=21us
avg=48us
max=241us
deadline=1000us
misses=0

Run these tests:

text
A. Only GPIO input.
B. GPIO + EC25 UART traffic.
C. GPIO + MQTT + verbose logging.
D. GPIO + NVS commit / test Flash write.

6. What to read or try next

  • ESP-IDF FreeRTOS SMP scheduler, critical sections, ESP Timer, GPTimer and apptrace.
  • CMSIS-RTOS2 Thread Management, Mutex Management and Message Queue.
  • Response-time analysis for fixed-priority scheduling.
  • SystemView/Tracealyzer for visual analysis of ISR/task scheduling.

Overall review of lessons 41–50

Across these 10 lessons, we have moved from internal state architecture to production reliability:

text
FSM
→ binary protocol
→ fuzzing
→ fault injection
→ postmortem observability
→ reproducible release
→ Secure Boot / Flash Encryption
→ unique device identity
→ secure remote commands
→ latency budget

The series' key idea: a reliable embedded system is built on more than drivers and peripheral init. It requires explicit states, bounded buffers, testable protocols, secure updates, observability, reproducible releases and measured timing budgets.

A DMA half-buffer holds 128 samples and the aggregate sample rate is 32 ksample/s. What interval does the lesson use?

Criteria: Convert samples and rate consistently; do not infer a measured board result.

Exercise

A system averages 20% CPU load, yet its UART ring buffer overflows during Flash writes. Explain why the average is insufficient and name measurements that would guide a latency budget.

Self-check criteria: Use worst-case observed timing and buffer/deadline constraints. Distinguish intentional debounce from RTOS delay and do not claim an unperformed timing test.

Show the supplied answer

A long critical section or Flash erase/write can delay the consumer despite low average load. Measure maximum event-to-response delay and its interrupt, queue, scheduling and processing components; record histograms, deadline misses and drops under the specified GPIO/UART/MQTT/storage loads.