1. Today's topic

Configuration must be updated while real-time tasks continue reading it. For example:

text
phase thresholds
debounce/hysteresis
GPIO masks
EC25 timeouts
retry/backoff policy
MQTT keepalive
CAN/RS-485 parameters

The main idea: a reader must not see configuration that the writer is still modifying. The writer constructs a new immutable snapshot and publishes it with one atomic operation.

2. Why this matters

For the phase task, taking a mutex on every ADC block adds jitter. For EC25, it is important not to get a mixture of old and new at_timeout_ms, retry_limit and max_backoff_ms. For MQTT, hostname, port and topic prefix must not be read while they are partially updated.

3. Theory

Why atomic fields are insufficient

If every field is made atomic, a reader can still observe a mixture of generations:

text
threshold = generation 42
hysteresis = generation 41
debounce = generation 42

The entire snapshot must be published atomically.

Immutable snapshot

c
typedef struct {
    uint32_t schema;
    uint32_t generation;
    phase_config_t phase;
    modem_config_t modem;
    mqtt_config_t mqtt;
} app_config_snapshot_t;

After publication, the object is treated as const.

Double buffering and the lifetime problem

An A/B scheme solves publication, but not the lifetime of the old snapshot. A reader may have taken a pointer to A; the writer publishes B and starts rewriting A, causing a race.

Grace periods and hazard pointers

c
/* Editorial alternative: a queue copies runtime_config_t by value.
 * Only phase_task owns and reads local_cfg; no shared pointer is reclaimed.
 * Create a bounded queue of sizeof(runtime_config_t) before starting tasks.
 * Producers validate candidates and report a full queue instead of blocking forever.
 */
static void phase_apply_pending_config(QueueHandle_t queue,
                                       runtime_config_t *local_cfg)
{
    runtime_config_t candidate;
    /* At most one candidate per cycle: bounded fast-path work. */
    if (xQueueReceive(queue, &candidate, 0) == pdPASS &&
        runtime_config_validate(&candidate)) {
        *local_cfg = candidate;
    }
}
/* Call at a cycle boundary, then process the whole cycle using local_cfg.
 * Define generation/epoch acceptance separately; validation is not authorization.
 */

Linux RCU: lifetime and grace periods

Each reader publishes a hazard pointer: “I am using this snapshot now.” Reader pattern:

c
static const app_config_snapshot_t *config_read_acquire(config_reader_id_t reader)
{
    for (;;) {
        app_config_snapshot_t *snapshot = atomic_load_explicit(&s_active_config, memory_order_acquire);
        atomic_store_explicit(&s_hazards[reader], snapshot, memory_order_release);
        app_config_snapshot_t *verify = atomic_load_explicit(&s_active_config, memory_order_acquire);
        if (snapshot == verify) {
            return snapshot;
        }
        atomic_store_explicit(&s_hazards[reader], NULL, memory_order_release);
    }
}

Release:

c
static void config_read_release(config_reader_id_t reader)
{
    atomic_store_explicit(&s_hazards[reader], NULL, memory_order_release);
}

Keep the read section short

Do not retain a shared snapshot across a network operation or vTaskDelay(). It is better to copy a small subconfiguration and release the shared snapshot immediately.

c
phase_config_t local;
const app_config_snapshot_t *cfg = config_read_acquire(CONFIG_READER_PHASE);
local = cfg->phase;
uint32_t generation = cfg->generation;
config_read_release(CONFIG_READER_PHASE);
phase_process(samples, &local, generation);

Persistent and runtime configuration

A persistent update and runtime publication are different transaction boundaries. For durable remote configuration:

text
parse -> validate -> write inactive persistent slot -> commit -> read-back -> publish runtime snapshot -> ACK backend

For risky network settings, a trial mode is useful: apply them at runtime first, test connectivity, then persist them.

4. Common mistakes

  1. Modifying the active snapshot in place.
  2. Treating an atomic pointer as a complete RCU implementation.
  3. Reusing the inactive buffer immediately after a swap.
  4. Holding a hazard across a network operation.
  5. Keeping a pointer to a temporary string inside a snapshot.
  6. Updating individual atomics instead of the whole configuration.
  7. Publishing a candidate before validation.
  8. Validating configuration in the fast path.
  9. Forgetting the generation.
  10. Acknowledging the backend before the persistent commit.
  11. Persisting risky network configuration without a trial.
  12. Using RCU where an owner-task local copy is sufficient.
  13. Reading a complex snapshot from an ISR.

5. Practical task

Create runtime configuration for phase and EC25:

c
typedef struct {
    uint16_t on_mv;
    uint16_t off_mv;
    uint16_t debounce_ms;
} phase_channel_config_t;
typedef struct {
    phase_channel_config_t red;
    phase_channel_config_t yellow;
    phase_channel_config_t green;
} phase_runtime_config_t;
typedef struct {
    uint32_t at_timeout_ms;
    uint32_t registration_timeout_ms;
    uint32_t backoff_ms;
    uint32_t max_backoff_ms;
    uint8_t retry_limit;
} modem_runtime_config_t;
typedef struct {
    uint32_t schema;
    uint32_t generation;
    phase_runtime_config_t phase;
    modem_runtime_config_t modem;
} runtime_config_t;

Implement a validator:

c
static bool runtime_config_validate(const runtime_config_t *cfg)
{
    if (cfg == NULL || cfg->schema != 1U) return false;
    if (cfg->phase.red.off_mv >= cfg->phase.red.on_mv) return false;
    if (cfg->modem.backoff_ms > cfg->modem.max_backoff_ms) return false;
    if (cfg->modem.retry_limit > 10U) return false;
    return true;
}

Define two reader IDs: PHASE and MODEM. Write a host stress test: the writer changes generation while readers continuously check invariants.

6. What to try next

  • Try owner-task local configuration first: for phase_task, sending a new subconfiguration in a message is often simpler.
  • For an ISR, publish only small precomputed atomic scalars.
  • For rapid consecutive updates, move from double buffering to triple buffering.
After atomically publishing snapshot B, may the writer immediately overwrite old snapshot A?

Exercise

A phase reader takes A, then a writer publishes B. Describe the race if the writer rewrites A immediately. Give the simpler owner-task alternative suggested by the lesson, without claiming that the shown atomic operations form a complete portable RCU implementation.

Self-check criteria: Separate publication from lifetime; describe the reader/writer race and a task-owned local-copy alternative.

Show the supplied answer

The reader can observe A while it is being rewritten, so its configuration is no longer an immutable coherent snapshot. The writer must wait for the required lifetime/reclamation condition, or send a validated subconfiguration message to the phase owner task, which updates and reads its own local copy at a defined event boundary.