1. Today's topic
We test a parser not only with manual tests, but with millions of random and semi-random inputs:
random bytes
corrupted lengths
truncated headers
incomplete COBS
a missing delimiter
thousands of zero-length TLV records
arbitrary chunking of the UART streamTools:
libFuzzer
AddressSanitizer
UndefinedBehaviorSanitizer
seed corpus
fuzz invariants
crash minimization
regression corpusThe main idea: the parser must be independent of the MCU, RTOS and HAL so it can run on a host under sanitizers millions of times per second. Timing, DMA, ISRs and integration are then checked on the board.
2. Why this matters
For a HIL protocol, an ordinary test checks PING, BAD_CRC and BAD_VERSION. A fuzzer finds unusual combinations:
header_length = 63
payload_length = UINT32_MAX - 40
TLV length = 65535
the actual frame is 17 bytes long
the next frame is joined to the previous oneFor an EC25 AT parser, fuzzing helps check:
a long +QMTRECV
embedded NUL
CR without LF
a URC inside a response
a late OK after timeout
a payload larger than maximum3. Theory
Unit tests versus fuzzing
A unit test checks a known input. A fuzzer generates new inputs, tracks coverage and keeps inputs that reach new branches. BAD_FORMAT is not a test failure; the following are:
out-of-bounds read/write
use-after-free
integer overflow
undefined behavior
assert invariant failure
timeout/hangSanitizers
ASan catches:
heap/stack/global overflow
use-after-free
double free
invalid freeUBSan catches:
signed overflow
bad shift
misaligned access
null dereferenceBuilding the fuzz target:
clang -std=c17 -O1 -g -fno-omit-frame-pointer \
-fno-sanitize-recover=undefined \
-fsanitize=fuzzer,address,undefined \
protocol/*.c tests/fuzz/protocol_stream_fuzz.c \
-Iprotocol -o build/protocol_stream_fuzzChunking invariance
The same stream must produce the same result regardless of feed() boundaries:
feed(all bytes once) == feed(bytes by random chunks)This is critical for UART/TCP. An example invariant:
fuzz_result_t one = parse_as_one_chunk(data, size);
fuzz_result_t chunks = parse_as_chunks(data, size);
if (one.frame_count != chunks.frame_count || one.digest != chunks.digest) {
__builtin_trap();
}Separate fuzz targets
CRC makes it harder for a fuzzer to reach deeper code, so separate targets are needed:
stream fuzz: COBS + CRC + resync
frame fuzz: decoded frame parser
TLV fuzz: payload parser directly
message fuzz: a specific schema
FSM fuzz: an event sequence
AT fuzz: stream/line/transaction layers4. Common mistakes
- Fuzzing the entire firmware at once.
- Not enabling sanitizers.
- Failing to reset a global parser between inputs.
- Treating malformed input as a test failure.
- Printing every input.
- Allowing CRC to block all deeper parser branches.
- Not limiting input size.
- Checking only for crashes rather than invariants.
- Not saving the input that caused a crash.
- Failing to add a fixed crash case to the regression corpus.
5. A practical task for 30–60 minutes
Create this structure:
tests/fuzz/
protocol_stream_fuzz.c
corpus/protocol_stream/
artifacts/
regressions/Add a fuzz target for proto_rx_feed():
input
→ feed as a whole
→ feed in chunks
→ compare the digest of accepted framesCreate seeds using the production encoder:
a valid PING
two consecutive PING frames
a frame with an optional TLVRun:
./build-fuzz/protocol_stream_fuzz \
tests/fuzz/corpus/protocol_stream \
-artifact_prefix=tests/fuzz/artifacts/ \
-max_len=4096 \
-timeout=2 \
-max_total_time=60 \
-print_final_stats=1Check the test infrastructure: temporarily introduce an off-by-one error in a buffer, wait for an ASan report, save the crash input as a regression case, then restore the fix.
6. What to read or try next
- LLVM libFuzzer: corpora, dictionaries, minimisation and parallel jobs.
- Clang ASan and UBSan.
- AFL++ for long campaigns.
- ESP-IDF host tests and Unity target tests.
Criteria: Separate an expected rejection from a safety or consistency failure.
Exercise
Design a chunking-invariance check for proto_rx_feed() and explain what to do with a minimized input that exposes a defect.
Self-check criteria: Use identical bytes, clean parser state, equivalent output comparison and a retained regression input; do not treat ordinary malformed-input rejection as a crash.
Show the supplied answer
Reset independent parser instances, feed the same byte stream once as a whole and once in chunks, then compare the digest of accepted frames. Save the minimized failing input in the regression corpus, fix the defect and rerun it with sanitizers. Avoid state leaking between inputs.