ADR-017: The firmware runs a stepped loop over a bounded monotonic delta
Date: 2026-08-07
Status: Accepted
Context
apps/firmware was a bounded probe. pet_firmware_run created a world, applied one caller-supplied
elapsed time, took a snapshot, reset the world, and returned. It proved that the Core builds and
links for esp32 and nothing else, because a lifecycle that ends cannot show that a pet lives.
The device has a monotonic counter and no wall clock. esp_timer_get_time returns microseconds since
boot, and only differences between two readings mean anything. The Core reads no clock at all, so
whatever reaches pet_world_update is measured and bounded by the host.
The Core already refuses an elapsed time above PET_TIME_DELTA_MAX_MS, which is one day. That bound
protects the world from a suspended host, and it is far too large to protect a device loop. A stall,
a debugger pause, a flash operation, or a counter anomaly can produce a gap that the Core would
accept and that would age the pet by time nobody lived through.
The loop is also where a device fails in ways a host build cannot show. A loop that never yields starves the idle task and its watchdog. A loop that returns leaves a board that boots and then does nothing. Neither is visible to a compiler.
Decision
The firmware reads time through a seam, bounds every delta before the Core sees it, and runs one step function forever.
PetFirmwareClock carries a monotonic now_ms and a yield_ms, both supplied by the platform. The
application calls neither an SDK function nor a standard library clock. platforms/esp_idf
implements the seam over esp_timer_get_time and a sleep the SDK turns into a task delay, and a host
test implements it over a supplied sequence of values.
pet_firmware_bounded_delta is pure. It takes the previous and the current reading and returns both
the measured elapsed time and what may be delivered from it. A reading that stalled or went backwards
gives zero, so a corrected counter cannot become an enormous unsigned wrap. A gap longer than
PET_FIRMWARE_DELTA_MAX_MS gives that bound, and the skipped time is dropped rather than replayed.
Each step then measures from the reading it actually took, so the loop resumes on the real clock
instead of chasing a debt.
The bound is 1000 milliseconds. The loop runs a 40 millisecond iteration, of which a full-frame panel transfer is expected to cost about 13, so the bound is about twenty-five ordinary iterations. That is wide enough that pacing jitter, a slow transfer, and SDK work never clamp, and narrow enough that a pause of any real length cannot jump the world forward. It sits far inside the one-day Core bound, so every bounded delta is a delta the Core accepts.
A clamp is reported as time-clamped with both durations. Time the world was denied is a fact about
the run, and a silent clamp would make the pet’s age disagree with the transcript for no stated
reason.
The lifecycle has three parts. pet_firmware_runtime_init creates the world once and records the
reading the first step measures from. pet_firmware_runtime_step advances the world to one reading
and derives one snapshot. pet_firmware_runtime_run emits runtime-ready once, then steps and
yields forever and does not return. The device driver and the host tests share the step, so what the
gate asserts is what the board runs.
Consequences
The world outlives the iteration. State is no longer created and discarded per call, which is what makes activity progression and, later, input and presentation possible at all.
The firmware allocates nothing. The runtime is caller-owned storage, the world lives inside it, and
app_main holds it as an automatic. Nothing about the loop introduces a heap dependency.
The hosted gate can drive the runtime deterministically. A test supplies readings and asserts the resulting world and snapshot, including a clamp, a stalled clock, and a clock that went backwards, none of which a device would produce on demand.
A step that the Core refuses is reported once and the loop continues. A condition the loop cannot clear would otherwise repeat its line thirty times a second and bury the transcript that the soak observation depends on, so the fault is reported when it appears and again once it has cleared. The firmware still resets nothing and silences nothing, which is what architecture section 8.12 requires.
The pacing is a usleep rather than a vTaskDelay. The ESP-IDF entry compiles as strict ISO C99,
and the FreeRTOS headers require GNU C, so including them would mean either a second dialect inside
one component or a weaker standard for all of it. The SDK turns a sleep of at least one scheduler
tick into a task delay, which was confirmed in the linked image rather than assumed, so the idle task
runs and its watchdog is fed.
The iteration is 40 milliseconds rather than the 33 this decision first recorded, and the pacing call
compensates for its own scheduler. usleep guarantees at least the requested time and wakes on a
tick boundary, so a request of exactly N ticks wakes just short of the target and costs N plus one.
The first device run measured 40 milliseconds against an announced 33, and moving the constant to 40
produced 50, which is what identified the extra tick as the cause rather than the value. BUG-001
records both measurements. The platform seam now subtracts one tick before it sleeps, with the tick
derived from the pinned CONFIG_FREERTOS_HZ, so the announced pace and the run pace are the same
number. The compensation belongs to the platform because the behaviour is the platform’s, and the
application still asks for the pace it wants.
The bound is a number this release chose against an expected iteration cost. The soak observation in
PBI-112 measures the real pace, and a bound that proves wrong is changed there with its reason
rather than defended here.
Alternatives considered
Read the SDK clock at the call site. Rejected because the elapsed time policy would then be untestable on the host, and the one piece of the loop with real decisions in it would only ever run on hardware.
Let the platform measure the delta and hand it to the application. Rejected because the bound, the backwards reading, and the clamp report are product-adjacent decisions that belong to the host application. A platform that measured them would own a policy every future platform would restate.
Reject a delta above the bound instead of clamping it. Rejected because the Core would refuse the update and the step would fail, so a single long pause would turn into a fault. A long pause is ordinary on a device that is being flashed or debugged, and resuming is the correct response.
Replay the skipped time in bounded chunks until the clock is caught up. Rejected because the work of one iteration would become unbounded in a way that the watchdog can see, and because the pet would then live through time that the loop did not observe. Dropping the difference and saying so is the smaller lie and the visible one.
Keep pet_firmware_run beside the stepped form. Rejected because after the loop exists nothing
calls it and no test drives it. A one-shot lifecycle that initialises and resets a world per call
also contradicts the rule the loop depends on, which is that the world outlives the iteration.
Put the loop in app_main and keep the application a step function. Rejected because the loop
holds the pacing, the readiness event, and the fault policy, and a platform that owned those would
own the parts of the runtime a host test can reach. The application owns the lifecycle, which is what
the general architecture already assigns to a host.
Drive the loop from a FreeRTOS timer or a dedicated task per concern. Rejected because it buys concurrency the release has no use for and pays in SDK coupling and in a shape no host test can run. One loop with one bounded step is enough for a pet, a panel, and two buttons.
References
- Architecture for v0.5.0, sections 6.2, 8.1, and 8.12
- ADR-002, the explicit time contract this bounds
- ADR-012, the wall clock the device does not have
- Epics,
E32 - PBIs,
PBI-110andPBI-111