1. Today's topic
Configuration must be updated while real-time tasks continue reading it. For example:
phase thresholds
debounce/hysteresis
GPIO masks
EC25 timeouts
retry/backoff policy
MQTT keepalive
CAN/RS-485 parametersThe main idea: a reader must not see configuration that the writer is still modifying. The writer constructs a new immutable snapshot and publishes it with one atomic operation.
2. Why this matters
For the phase task, taking a mutex on every ADC block adds jitter. For EC25, it is important not to get a mixture of old and new at_timeout_ms, retry_limit and max_backoff_ms. For MQTT, hostname, port and topic prefix must not be read while they are partially updated.
3. Theory
Why atomic fields are insufficient
If every field is made atomic, a reader can still observe a mixture of generations:
threshold = generation 42
hysteresis = generation 41
debounce = generation 42The entire snapshot must be published atomically.
Immutable snapshot
typedef struct {
uint32_t schema;
uint32_t generation;
phase_config_t phase;
modem_config_t modem;
mqtt_config_t mqtt;
} app_config_snapshot_t;After publication, the object is treated as const.
Double buffering and the lifetime problem
An A/B scheme solves publication, but not the lifetime of the old snapshot. A reader may have taken a pointer to A; the writer publishes B and starts rewriting A, causing a race.
Grace periods and hazard pointers
/* Editorial alternative: a queue copies runtime_config_t by value.
* Only phase_task owns and reads local_cfg; no shared pointer is reclaimed.
* Create a bounded queue of sizeof(runtime_config_t) before starting tasks.
* Producers validate candidates and report a full queue instead of blocking forever.
*/
static void phase_apply_pending_config(QueueHandle_t queue,
runtime_config_t *local_cfg)
{
runtime_config_t candidate;
/* At most one candidate per cycle: bounded fast-path work. */
if (xQueueReceive(queue, &candidate, 0) == pdPASS &&
runtime_config_validate(&candidate)) {
*local_cfg = candidate;
}
}
/* Call at a cycle boundary, then process the whole cycle using local_cfg.
* Define generation/epoch acceptance separately; validation is not authorization.
*/Linux RCU: lifetime and grace periods
Each reader publishes a hazard pointer: “I am using this snapshot now.” Reader pattern:
static const app_config_snapshot_t *config_read_acquire(config_reader_id_t reader)
{
for (;;) {
app_config_snapshot_t *snapshot = atomic_load_explicit(&s_active_config, memory_order_acquire);
atomic_store_explicit(&s_hazards[reader], snapshot, memory_order_release);
app_config_snapshot_t *verify = atomic_load_explicit(&s_active_config, memory_order_acquire);
if (snapshot == verify) {
return snapshot;
}
atomic_store_explicit(&s_hazards[reader], NULL, memory_order_release);
}
}Release:
static void config_read_release(config_reader_id_t reader)
{
atomic_store_explicit(&s_hazards[reader], NULL, memory_order_release);
}Keep the read section short
Do not retain a shared snapshot across a network operation or vTaskDelay(). It is better to copy a small subconfiguration and release the shared snapshot immediately.
phase_config_t local;
const app_config_snapshot_t *cfg = config_read_acquire(CONFIG_READER_PHASE);
local = cfg->phase;
uint32_t generation = cfg->generation;
config_read_release(CONFIG_READER_PHASE);
phase_process(samples, &local, generation);Persistent and runtime configuration
A persistent update and runtime publication are different transaction boundaries. For durable remote configuration:
parse -> validate -> write inactive persistent slot -> commit -> read-back -> publish runtime snapshot -> ACK backendFor risky network settings, a trial mode is useful: apply them at runtime first, test connectivity, then persist them.
4. Common mistakes
- Modifying the active snapshot in place.
- Treating an atomic pointer as a complete RCU implementation.
- Reusing the inactive buffer immediately after a swap.
- Holding a hazard across a network operation.
- Keeping a pointer to a temporary string inside a snapshot.
- Updating individual atomics instead of the whole configuration.
- Publishing a candidate before validation.
- Validating configuration in the fast path.
- Forgetting the generation.
- Acknowledging the backend before the persistent commit.
- Persisting risky network configuration without a trial.
- Using RCU where an owner-task local copy is sufficient.
- Reading a complex snapshot from an ISR.
5. Practical task
Create runtime configuration for phase and EC25:
typedef struct {
uint16_t on_mv;
uint16_t off_mv;
uint16_t debounce_ms;
} phase_channel_config_t;
typedef struct {
phase_channel_config_t red;
phase_channel_config_t yellow;
phase_channel_config_t green;
} phase_runtime_config_t;
typedef struct {
uint32_t at_timeout_ms;
uint32_t registration_timeout_ms;
uint32_t backoff_ms;
uint32_t max_backoff_ms;
uint8_t retry_limit;
} modem_runtime_config_t;
typedef struct {
uint32_t schema;
uint32_t generation;
phase_runtime_config_t phase;
modem_runtime_config_t modem;
} runtime_config_t;Implement a validator:
static bool runtime_config_validate(const runtime_config_t *cfg)
{
if (cfg == NULL || cfg->schema != 1U) return false;
if (cfg->phase.red.off_mv >= cfg->phase.red.on_mv) return false;
if (cfg->modem.backoff_ms > cfg->modem.max_backoff_ms) return false;
if (cfg->modem.retry_limit > 10U) return false;
return true;
}Define two reader IDs: PHASE and MODEM. Write a host stress test: the writer changes generation while readers continuously check invariants.
6. What to try next
- Try owner-task local configuration first: for phase_task, sending a new subconfiguration in a message is often simpler.
- For an ISR, publish only small precomputed atomic scalars.
- For rapid consecutive updates, move from double buffering to triple buffering.
Exercise
A phase reader takes A, then a writer publishes B. Describe the race if the writer rewrites A immediately. Give the simpler owner-task alternative suggested by the lesson, without claiming that the shown atomic operations form a complete portable RCU implementation.
Self-check criteria: Separate publication from lifetime; describe the reader/writer race and a task-owned local-copy alternative.
Show the supplied answer
The reader can observe A while it is being rewritten, so its configuration is no longer an immutable coherent snapshot. The writer must wait for the required lifetime/reclamation condition, or send a validated subconfiguration message to the phase owner task, which updates and reads its own local copy at a defined event boundary.