ADR-017: The firmware runs a stepped loop over a bounded monotonic delta

Date: 2026-08-07

Status: Accepted

Context

apps/firmware was a bounded probe. pet_firmware_run created a world, applied one caller-supplied elapsed time, took a snapshot, reset the world, and returned. It proved that the Core builds and links for esp32 and nothing else, because a lifecycle that ends cannot show that a pet lives.

The device has a monotonic counter and no wall clock. esp_timer_get_time returns microseconds since boot, and only differences between two readings mean anything. The Core reads no clock at all, so whatever reaches pet_world_update is measured and bounded by the host.

The Core already refuses an elapsed time above PET_TIME_DELTA_MAX_MS, which is one day. That bound protects the world from a suspended host, and it is far too large to protect a device loop. A stall, a debugger pause, a flash operation, or a counter anomaly can produce a gap that the Core would accept and that would age the pet by time nobody lived through.

The loop is also where a device fails in ways a host build cannot show. A loop that never yields starves the idle task and its watchdog. A loop that returns leaves a board that boots and then does nothing. Neither is visible to a compiler.

Decision

The firmware reads time through a seam, bounds every delta before the Core sees it, and runs one step function forever.

PetFirmwareClock carries a monotonic now_ms and a yield_ms, both supplied by the platform. The application calls neither an SDK function nor a standard library clock. platforms/esp_idf implements the seam over esp_timer_get_time and a sleep the SDK turns into a task delay, and a host test implements it over a supplied sequence of values.

pet_firmware_bounded_delta is pure. It takes the previous and the current reading and returns both the measured elapsed time and what may be delivered from it. A reading that stalled or went backwards gives zero, so a corrected counter cannot become an enormous unsigned wrap. A gap longer than PET_FIRMWARE_DELTA_MAX_MS gives that bound, and the skipped time is dropped rather than replayed. Each step then measures from the reading it actually took, so the loop resumes on the real clock instead of chasing a debt.

The bound is 1000 milliseconds. The loop runs a 40 millisecond iteration, of which a full-frame panel transfer is expected to cost about 13, so the bound is about twenty-five ordinary iterations. That is wide enough that pacing jitter, a slow transfer, and SDK work never clamp, and narrow enough that a pause of any real length cannot jump the world forward. It sits far inside the one-day Core bound, so every bounded delta is a delta the Core accepts.

A clamp is reported as time-clamped with both durations. Time the world was denied is a fact about the run, and a silent clamp would make the pet’s age disagree with the transcript for no stated reason.

The lifecycle has three parts. pet_firmware_runtime_init creates the world once and records the reading the first step measures from. pet_firmware_runtime_step advances the world to one reading and derives one snapshot. pet_firmware_runtime_run emits runtime-ready once, then steps and yields forever and does not return. The device driver and the host tests share the step, so what the gate asserts is what the board runs.

Consequences

The world outlives the iteration. State is no longer created and discarded per call, which is what makes activity progression and, later, input and presentation possible at all.

The firmware allocates nothing. The runtime is caller-owned storage, the world lives inside it, and app_main holds it as an automatic. Nothing about the loop introduces a heap dependency.

The hosted gate can drive the runtime deterministically. A test supplies readings and asserts the resulting world and snapshot, including a clamp, a stalled clock, and a clock that went backwards, none of which a device would produce on demand.

A step that the Core refuses is reported once and the loop continues. A condition the loop cannot clear would otherwise repeat its line thirty times a second and bury the transcript that the soak observation depends on, so the fault is reported when it appears and again once it has cleared. The firmware still resets nothing and silences nothing, which is what architecture section 8.12 requires.

The pacing is a usleep rather than a vTaskDelay. The ESP-IDF entry compiles as strict ISO C99, and the FreeRTOS headers require GNU C, so including them would mean either a second dialect inside one component or a weaker standard for all of it. The SDK turns a sleep of at least one scheduler tick into a task delay, which was confirmed in the linked image rather than assumed, so the idle task runs and its watchdog is fed.

The iteration is 40 milliseconds rather than the 33 this decision first recorded, and the pacing call compensates for its own scheduler. usleep guarantees at least the requested time and wakes on a tick boundary, so a request of exactly N ticks wakes just short of the target and costs N plus one. The first device run measured 40 milliseconds against an announced 33, and moving the constant to 40 produced 50, which is what identified the extra tick as the cause rather than the value. BUG-001 records both measurements. The platform seam now subtracts one tick before it sleeps, with the tick derived from the pinned CONFIG_FREERTOS_HZ, so the announced pace and the run pace are the same number. The compensation belongs to the platform because the behaviour is the platform’s, and the application still asks for the pace it wants.

The bound is a number this release chose against an expected iteration cost. The soak observation in PBI-112 measures the real pace, and a bound that proves wrong is changed there with its reason rather than defended here.

Alternatives considered

Read the SDK clock at the call site. Rejected because the elapsed time policy would then be untestable on the host, and the one piece of the loop with real decisions in it would only ever run on hardware.

Let the platform measure the delta and hand it to the application. Rejected because the bound, the backwards reading, and the clamp report are product-adjacent decisions that belong to the host application. A platform that measured them would own a policy every future platform would restate.

Reject a delta above the bound instead of clamping it. Rejected because the Core would refuse the update and the step would fail, so a single long pause would turn into a fault. A long pause is ordinary on a device that is being flashed or debugged, and resuming is the correct response.

Replay the skipped time in bounded chunks until the clock is caught up. Rejected because the work of one iteration would become unbounded in a way that the watchdog can see, and because the pet would then live through time that the loop did not observe. Dropping the difference and saying so is the smaller lie and the visible one.

Keep pet_firmware_run beside the stepped form. Rejected because after the loop exists nothing calls it and no test drives it. A one-shot lifecycle that initialises and resets a world per call also contradicts the rule the loop depends on, which is that the world outlives the iteration.

Put the loop in app_main and keep the application a step function. Rejected because the loop holds the pacing, the readiness event, and the fault policy, and a platform that owned those would own the parts of the runtime a host test can reach. The application owns the lifecycle, which is what the general architecture already assigns to a host.

Drive the loop from a FreeRTOS timer or a dedicated task per concern. Rejected because it buys concurrency the release has no use for and pays in SDK coupling and in a shape no host test can run. One loop with one bounded step is enough for a pet, a panel, and two buttons.

References