Skip to content

ADR 0008: Edge-local holding logic — partition-triggered safe-state sequencing at the unit runtime

Status: Accepted Date: 2026-06-16 Issue: #578 Related: ADR 0006 (refines its partition-gap context), ADR 0007 (extends its three-layer protection model), #566 (edge redundancy scoping)

Context

ADR 0007 established a three-layer protection model and was explicit about which layer survives a control-plane partition:

Layer Where it runs Survives partition Role
Device interlock (DO/AO ILCK) FB scan, edge Yes Force one output to one static safe value while a trip is true
Phase SFC guard (interlock: true) Procedural operator, control plane No Procedural safe-state sequencing — ordered multi-output response (vent first, then cut heat), parking the procedure in a resumable held step
AlarmDefinition.exceptionAction Alarm controller, control plane No Annunciation, audit, batch-level Hold/Stop/Abort

The gap #578 names is the middle row. A device interlock forces a single output to a single value — it cannot sequence. The only layer that can run a deliberate, ordered safe-state response (the thing a process actually needs: close the charge valve, then stop the agitator, then open the vent) is the phase SFC, and that executes in the procedural operator on the control plane (internal/controller/procedural/phase_controller.go instantiates the engine; each scan's READ/WRITE builtins cross HTTP to the runtime via pkg/stbridge). On a control-plane partition:

  • the edge FB scan keeps regulating to whatever setpoints were last commanded (the autonomy claim — accurate, ADR 0006), but
  • the phase SFC stalls mid-step at an arbitrary point, so the equipment is frozen wherever the sequence happened to be — mid-charge with the charge valve held open is frozen, not safe, and
  • no sequenced holding response can execute at the edge, because that logic lives in the control plane.

ADR 0006 made this precise and accepted it as the v1 posture: "Field outputs hold last value (or device fail-safe) during the gap, exactly as for a pod restart today" (ADR 0006 §Decision.3). The compliance traceability says the same — docs/compliance/isa88.md Clause 7.4 row: "Held outputs hold last value / device fail-safe during the gap." Frozen, not driven to a deliberate safe state.

The pod-restart half of that sentence was not true until #1776. A restarting runtime rebuilt every block empty and wrote a compile-time default on its first scan, so a hold this ADR had armed correctly was discarded by the same fault's own restart. See ADR 0080.

Competitive + standards baseline (researched 2026-06-12, #578):

  • DeltaV's shipped default on Batch Executive ↔ controller comm loss is the phase placed in Held via a watchdog, with Holding logic executing locally in the controller with full I/O access — deliberate safe-state actions, not frozen outputs (Emerson Batch Executive PDS). Rockwell FactoryTalk Batch documents the same watchdog → HELD philosophy. Note this is the non-redundant comm-loss path; it is orthogonal to controller redundancy (ADR 0006), which is a separate feature.
  • ISA-88 Part 1 deliberately does not mandate execution placement (Clause 6.6.3: recipe/equipment separation is logical, physical separation optional; Clause 6.6.4: equipment procedural element internals are out of scope). What it expects is the Clause 7.4 exception response: malfunction → Hold to a known safe state → Restart with recipe-defined restarting logic.

So the gap is not "move the phase engine to the edge" — that would put mid-phase durable state on the node ADR 0006 deliberately allows to die, and complicate the 21 CFR Part 11 audit trail. The gap is narrow and specific: the edge has no locally executable, sequenced holding response.

Five constraints from the existing code shape the option space:

  1. One writer owns the drivers. ADR 0007 rejected a separate watchdog goroutine forcing driver writes precisely because "two writers to one output address race each other; the FB scan already owns output ordering and the one-writer guarantee." Any edge holding response must respect this — it cannot become a second driver writer.
  2. SFC encapsulation rule. All phase state logic must be SFC charts; no bare flat ST (project convention). A sequenced hold response is phase state logic.
  3. The SFC engine is already standalone. pkg/sfc imports only the CRD types, the ST interpreter, and a clock; it accepts an injectable StepHandler and supports WithRestoredState. It does not today have a driver-bound execution path at the edge — its READ/WRITE resolve over HTTP in the control plane.
  4. Failover self-fence and a local hold would fight. ADR 0006's Failover mode self-fences on lease expiry — it stops FB output writes so a standby can take over (internal/adapter/lease.go fencedDriver drops writes). A holding response that writes during that window contradicts the fence and risks dual writers against the standby.
  5. Durable Held state lives in etcd, and the phase self-holds anyway. Phase state lives in Phase.Status (ADR 0006); the procedural operator's existing 60 s grace → self-Hold path (phase_controller.go) independently drives the phase to Held when it sees the runtime unreachable. Whatever the edge does must converge with that, not diverge from it.

Decision

The unit runtime gains an armed local-hold program: an SFC chart executed by an SFC engine embedded in the runtime when the edge's control-plane heartbeat watchdog fires. The hold chart drives a deliberate, sequenced safe state by writing the unit's own control-module tag space — the FB scan remains the sole I/O writer, so no new write path and no dual-writer is introduced. This makes ADR 0007's procedural safe-state layer partition-tolerant. Durable Held state continues to live in the control plane (ADR 0006); recovery is a normal ISA-88 Restart (Clause 7.4) — no silent resume. This refines ADR 0006's partition gap; it does not supersede it.

The scoped mechanism:

  1. What is armed. Two sources, in priority order:
  2. Per-phase (preferred): at phase start the procedural operator downloads the active phase's HoldingChart to the runtime as the currently-armed local-hold program — the same chart that runs the ISA-88 Holding state in the control plane, now also staged at the edge. This reproduces the DeltaV default (phase-specific holding logic) and is correct precisely because a mid-charge partition arms the charge phase's holding logic.
  3. Unit baseline (fallback): a new UnitSpec.safeStateChart (SFC), deployed with the unit's control program and always armed when no phase is active. This closes ADR 0007's idle-unit / manual-mode coverage gap for the sequenced case, the same way device interlocks closed it for the single-output case.

Edge-armable constraint. A chart is armable at the edge only if it READ/WRITEs the unit's own control-module tag space (no cross-unit, no control-plane-only data). A chart referencing out-of-edge-scope data cannot be staged; the runtime falls back to the unit baseline and the operator surfaces a warning. A lint/validation enforces this.

  1. Execution model. The runtime embeds the pkg/sfc engine with a StepHandler/tag binding whose READ/WRITE resolve directly against the local control-module tag space (the FB network's variable bindings) rather than over HTTP. The hold chart manipulates FB inputs (setpoints, valve commands) exactly as the phase SFC does today via pkg/stbridge; the FB scan regulates to them and remains the only writer to drivers. Device interlocks (ADR 0007) still sit underneath and can override even the hold chart's commanded outputs. The SFC encapsulation rule is honored — the hold response is an SFC chart.

  2. Trigger contract — a unified control-plane heartbeat. The local-hold trigger is "no control-plane heartbeat within the timeout," evaluated by a purely local timer at the edge (no network reach required to decide):

  3. In Failover mode the heartbeat is the existing lease renew (/api/v1/lease/renew, internal/adapter/lease.go) — reuse it, do not add a second watchdog.
  4. In Autonomy mode there is no edge watchdog today (the lease guard is inert). Add a lightweight liveness heartbeat the procedural/physical operator issues each reconcile (carrying "phase X is live") — no fencing, no lease semantics, just a timestamp the edge watchdog consumes.
  5. Timeout is the unit's availability.leaseDurationSeconds in Failover, and a new availability.holdGraceSeconds (default mirrors the control plane's 60 s phase self-hold grace) in Autonomy.

  6. Mode interaction (resolves constraint 4).

  7. Autonomy: on heartbeat loss past the grace window, the edge runs the armed hold chart to drive a deliberate safe state, then keeps holding (the FB scan continues regulating to the safe-state setpoints the chart established). No fence — the edge is the legitimate single writer.
  8. Failover: on lease expiry the edge runs the armed hold chart as a bounded safe-state sequence, then self-fences (stops writes) so a standby can take over cleanly. The standby re-bind already happens only after lease-expiry + margin (ADR 0006); the margin must be ≥ the hold sequence's bounded duration, so the partitioned node finishes its safe-state actions and fences before the standby writes — no dual writer. A node that is truly dead (not merely partitioned) cannot run the hold; that case falls back to ADR 0006's existing behavior: outputs at device fail-safe / last value until the standby re-binds and Restart drives them. What the re-bind itself writes is ADR 0082, and until then it was a compile-time default rather than a continuation. Local-hold strictly improves the alive-but-partitioned case; it does not regress the dead-node case.

  9. Reconciliation on reconnect. The runtime reports, via its status (MQTT status + a gRPC/HTTP field), that it self-held, which program it ran, and when. On reconnect the operator reads this and reflects Held on the Phase/Unit state machines. This converges with the control plane's own 60 s self-hold path (constraint 5): during a partition both sides independently move toward Held — etcd marks the phase Held, the edge holds outputs — and on reconnect they agree. Recovery requires an explicit ISA-88 Restart (Clause 7.4); there is no silent resume.

  10. Audit at the edge. Local-hold executes while the control-plane audit path is unreachable. The runtime emits hold-lifecycle events (armed, triggered, each safe-state action, self-fenced) onto the existing MQTT store-and-forward queue (internal/adapter/queue, /var/lib/dcs/runtime/queue/), replayed on reconnect and materialized into AuditRecord CRs by the audit bridge. This preserves the "audit must not block the control loop" rule and reuses the historian store-and-forward precedent rather than giving the runtime direct apiserver write access.

Onto the queue always, never around it (#1779). The first implementation of this sink tried a live publish and fell back to the queue, which is the right shape for telemetry and the wrong one here. Hold events are emitted on the goroutine driving the equipment to its safe state: the triggered event goes out ahead of the engine, and each safe-state action goes out from inside the step that performed it, which does not advance until its own emission returns. A live publish at QoS 1 waits for a PUBACK, and during a control-plane partition the broker's answer is exactly what nothing can promise. Measured on the bench, a hold that settles in 300ms left the field outputs at their pre-fault values for 10.3 seconds. The sink therefore appends to the queue and wakes the replay worker, which does the network part on its own goroutine. Bounding the publish was rejected as the fix, because a bound still spends the bound. The queue is also the more honest record: the event is durable before control moves on, where a live publish still in flight when the pod dies leaves nothing behind.

The bench rep ran in Autonomy, where the hold has no deadline and the delay costs 10.3 seconds of equipment sitting at its pre-fault setpoint. In Failover it costs more than that. The hold there is bounded at half the lease, clamped to between 5 and 30 seconds, and the stall sits ahead of the engine rather than inside it. A stall long enough to consume that budget would have the runtime self-fence having written nothing at all, and a standby would take over the frozen unit that the bounded hold exists to prevent.

The trail carries the order, and the timestamp never could (#1813). This decision is about a sequenced safe state, so the order of the actions is the content of the record and not a detail of it. The recorded timestamp cannot carry it. AuditRecordSpec.timestamp is a metav1.Time, which serialises at whole-second precision, while the hold engine scans every 100 ms, so the two WRITEs of a two-step safe state land on one second and tie. The gateway then sorted the trail on that value with an unstable sort, and two reads of one unchanged trail could disagree about which output was driven to its safe value first. Each event therefore carries an emission ordinal, assigned by the runtime under its own lock at the moment of emission rather than derived from a clock. That distinction matters here more than anywhere else in the product. The node is partitioned, its clock may step when the network returns, and the ordinal is unaffected by either. The ordinal rides the hold snapshot, so a pod that restarts mid-partition resumes numbering where it left off instead of renumbering over the events already queued. The published guarantee and the ordering every audit surface applies are recorded in 21 CFR Part 11 Compliance Traceability.

Alternatives Considered

  • Download a compiled safe-state FB network instead of an SFC chart (engine already at the edge — issue option 2). Rejected as the primary: a flat FB network cannot sequence (vent-then-cut-heat), which is the entire gap — ADR 0007 keeps sequencing in the SFC layer for exactly this reason, and the FB layer already provides the static single-output safe state via device interlocks. Re-encoding ordered sequences as FB logic would duplicate the SFC engine's job and violate the SFC encapsulation rule.
  • Move the whole phase/SFC engine to the edge. Rejected: it puts durable mid-phase state (active steps, fired transitions, variable values) on the node ADR 0006 deliberately allows to die, breaks the etcd-resident phase state that lets a running phase survive failover, and complicates the Part 11 audit trail. The fix is a fallback safe-state responder at the edge, not relocating sequencing ownership.
  • Unit-level static safe-state chart only (no per-phase arming). Rejected as the sole answer: it loses phase-specific safe states (mid-charge wants a different response than mid-heat), which is the DeltaV parity the issue is chasing. Kept as the baseline for the idle/no-phase case.
  • Status quo — frozen outputs + device interlocks only (ADR 0006 v1 posture). Rejected as terminal: a device interlock cannot sequence, and frozen-at-last-value is not a deliberate safe state. The diligence comparison against the DeltaV/Rockwell shipped default fails on exactly this point.
  • A separate watchdog goroutine forcing driver writes at the edge. Rejected for the same reason ADR 0007 rejected it for interlocks: two writers to one address race; the FB scan owns the one-writer guarantee. The embedded SFC engine writes the FB variable space, never the drivers directly.

Consequences

  • API / code surfaces that move:
  • UnitSpec gains safeStateChart (SFC, the unit baseline) and availability.holdGraceSeconds (Autonomy trigger timeout).
  • The unit runtime embeds the pkg/sfc engine with a driver-bound tag handler (internal/adapter/, cmd/unit-runtime/); a new edge watchdog consumes the heartbeat; new self-held status field on the adapter API.
  • The adapter API gains an arm-hold-program endpoint (download the active phase's HoldingChart) and, for Autonomy, a liveness heartbeat endpoint; Failover reuses /api/v1/lease/renew.
  • The procedural operator downloads the phase HoldingChart at phase start and reconciles edge-self-held → Held on reconnect (phase_controller.go).
  • The audit bridge consumes edge hold-lifecycle events from the MQTT queue and materializes AuditRecord CRs.
  • New metrics: hold armed/triggered counters, hold-sequence duration, edge-self-held gauge.
  • Compliance:
  • docs/compliance/isa88.md Clause 7.4 row (control equipment malfunction → Hold → Restart) is refined: the partition response becomes a deliberate sequenced safe state, not frozen-at-last-value. The "Held outputs hold last value" caveat narrows to the dead-node / no-armed-chart case. docs/library/alarms-and-interlocks.md's three-layer table gains a note that the procedural safe-state layer is now partition-tolerant at the edge.
  • This is a candidate execution vehicle for the existing ProcessException safeState.structuredText row (currently "Partial — validated and logged, not executed against equipment"): an armed unit baseline chart is where a process-exception safe state could actually run. Noted for the implementation epic, not decided here.
  • 21 CFR Part 11: the edge-buffered, replayed audit trail records the hold actions taken while the control plane was unreachable.
  • Docs that update when this ships: docs/ha-failure-modes.md (partition row: deliberate safe-state, not frozen), docs/architecture.md (autonomy section), docs/adr/0007 three-layer narrative cross-link.
  • Marketing posture: now that this has shipped (epic #597, merged), the accurate partition claim is "watchdog-triggered local holding logic drives a deliberate safe state at the edge during a control-plane partition" — the DeltaV/Rockwell default reproduced. The pre-ship claim ("autonomy + restart-replay + supervised Hold, outputs frozen at last value") now applies only to the dead-node / no-armed-chart case. Never claim bumpless redundancy (that remains explicitly rejected, ADR 0006).
  • Default behavior unchanged: with no safeStateChart and no phase HoldingChart, the edge behaves exactly as today (frozen outputs). The feature is additive and opt-in per unit/phase.
  • Reversibility: moderate now that this has shipped — UnitSpec.safeStateChart is a public CRD contract and the heartbeat is a runtime API surface. (It was high before implementation, when this was still a design posture.)
  • Follow-ups (implementation sub-issues): embed driver-bound SFC engine at the edge; UnitSpec.safeStateChart + holdGraceSeconds API; Autonomy liveness heartbeat + edge watchdog; per-phase HoldingChart arming at phase start; Failover hold-then-fence ordering + margin guard; edge-self-held status + reconnect reconciliation to Held; edge audit buffering of hold events; edge-armable chart validation/lint; docs + compliance updates.

Amendment (2026-08-10, #1406): edge-armable covers capabilities, not only data

The edge-armable constraint in §Decision.1 was written about data — the chart may address only the unit's own control-module tag space. It said nothing about capabilities, and the two fail the same way. Eleven of the ST dialect's builtins cannot execute at the edge at all, because each one needs something the partition has taken away: an operator (PROMPT, PROMPT_CHOICE, PROMPT_VALUE), the gateway an external system delivers a measurement through (AWAIT_RESULT, ADR 0055), the apiserver (MODE), the phase state machine the control plane drives (COMMAND), control-plane state about where the action chart stopped (STEP_ACTIVE), or a network session to another machine (CALL_SERVICE and the three MTP builtins, ADR 0045).

A chart calling one of them passed every gate, ran correctly on the control-plane path every time it was exercised, and errored mid-sequence only during the partition it was armed for. That is the worst moment to discover it, and it was discovered as an ST runtime error inside a hold, with no operator and no control plane to report it to.

The decision is unchanged and its scope is stated fully. Such a chart is not armable, which is the outcome §Decision.1 already defines: the runtime refuses to stage it, the phase falls back to the unit baseline safe-state chart, and the operator is told — while the control plane is still up to be told through. Authoring one remains legitimate, because a holding chart that prompts the operator, or commands a PEA service to hold, is exactly right for the ordinary control-plane hold it will normally run in. What changes is that its edge cover is one posture coarser, and that fact now arrives at arming time rather than during the partition.

ValidateEdgeArmableBuiltins (api/procedural/v1alpha1/edge_armable.go) is the static check, the runtime applies it when a chart is staged (internal/adapter/hold.go), and make lint-edge-armable enforces it over the example corpus alongside the tag-space half. Because the two halves must not drift, TestEdgeUnavailableBuiltinsMatchHoldRuntime runs every builtin the ST package implements through the real hold environment and fails if the list and the runtime disagree in either direction.

Amendment (2026-08-24, #1804): a recovered snapshot is inherited, and what it says is about a program

§Decision.5 says recovery requires an explicit ISA-88 Restart and that there is no silent resume. That is right about a hold this runtime ran. It was applied to a hold this runtime only found, and those are not the same claim.

HoldController.Recover restores the self-held flag from disk on startup, so a pod that restarted mid-partition reports the truth the moment it comes up. What it cannot know is whether the partition that raised the hold is still on. If it is, the watchdog or the lease re-triggers RunHold and this process owns the hold. If it is not — the pod came up into a working control plane, carrying a snapshot of a hold that ended — then the road §Decision.5 names has no traveller: POST /api/v1/hold/release rides a Restart, and the batch that produced the snapshot is finished. There is nothing to restart.

Measured on the bench during the drill-9 sitting for #1793. bench-loop-01's runtime pod came up at 17:22:49 carrying drill 3's 16:51 self-fence, three unrelated batches then ran to Complete on it, and dcs_runtime_self_held read 1 through all three — 361 scrapes at 1 Hz, no failures — while the loop regulated at 59.96–60.06 % PV and dcs_runtime_fenced stayed 0.

Two things follow, and they are separate defects with a shared root.

The runtime retires what it inherited. The flag now records whether it was raised here or found on disk. A hold this process ran is released by a Restart and by nothing else, unchanged. A hold it inherited is retired by RetireInheritedHold on evidence that the control plane is running this runtime again: a lease grant in Failover, a heartbeat in Autonomy. Either one says the road a Restart travels is open, so a claim still standing on it is about work that is over. The released audit event names which road was taken — control-plane-restart or inherited-snapshot-superseded — so the trail can tell an operator's recovery from a retirement.

The cost is stated rather than argued away. A pod that restarts mid-partition, whose partition then heals in the seconds before the control plane polls GET /api/v1/status, loses the reflection. It does not lose the Hold, which the control plane's own 60 s self-hold path lands on a phase whose runtime was out of reach that long (§Decision.5's convergence, from the other side), and it does not lose the record, which is §Decision.6's hold-event trail. What shipped instead was a signal pinned at 1 for the life of a pod, which fails towards alarm rather than towards all-clear and is useless in both directions — #1646's rule in its quieter shape.

A reflection is about a program, not about a flag. reconcileEdgeSelfHold drives the phase to Held without running its HoldingChart, on the reasoning that the edge already ran it. That reasoning belongs to the hold's own phase. Taken at face value, a self-held flag left over from other work drove a healthy running phase to Held having established no safe state at all, told the operator its runtime self-held during a partition that was not happening, and named a phase from a batch half an hour gone. The consequence is not confined to Autonomy: shouldPollEdgeHold fires on any phase that is Held or has lost contact with its runtime, in either availability mode.

edgeSelfHoldIsThisPhases is the discriminator, and it takes two readings because the program name settles only one case. A phase-scoped program names its phase outright and phase names are unique per batch. The unit baseline names nothing — a partition during a phase with no HoldingChart legitimately runs it — so time settles that one: a hold triggered before this phase started running is not this phase's hold. Both readings abstain rather than refuse where they cannot tell, which is §Decision.5's direction: reflecting a hold that turns out not to be ours costs a Restart, and missing one costs the reflection this ADR exists for.

TestReconcile_EdgeSelfHeld_IgnoresAnotherPhasesHold drives the bench snapshot in front of a later batch's phase and fails against the old shape; TestReconcile_EdgeSelfHeld_ReflectsABaselineHoldFromThisRun is what keeps the guard from being a blanket refusal. On the runtime side, TestRetireInheritedHold_LeavesAHoldThisProcessRan and TestRetireInheritedHold_RunHoldReclaimsARecoveredSnapshot hold the line the retirement road must not cross.

Amendment (2026-08-24, #1806): the baseline is armed by the unit, not by the transport that carries liveness

§Decision.1 says the unit baseline is "always armed when no phase is active", and §Decision.4 says separately that a Failover unit runs "the armed hold chart" as a bounded sequence on lease expiry. Both are about the unit. The implementation put the arming inside the Autonomy liveness heartbeat loop — §Decision.3's transport — and reconcileAvailability stops that loop in Failover mode, so UnitSpec.safeStateChart reached a Failover runtime on no road at all. The two paragraphs above were describing a posture the product had for one availability mode out of two.

What the gap costs depends on the phase. A running phase that declares a holdingChart arms it in preference and covers the window by accident, for as long as it runs. A phase that declares none — which is the ordinary case, and the case the bench drill for the unit chart exists to measure — had nothing armed, and the bounded pre-fence hold §Decision.4 promises ran a program that did not exist. The posture was frozen outputs, which is precisely what §Context rejects as terminal.

Measured on the bench, whose one unit is Failover with a safeStateChart declared. dcs_runtime_hold_armed read 0 at every drill rep that took the reading against a runtime no phase had armed, and the diagnosis the harness offers an operator — check the physical-operator log for arming unit baseline hold chart failed — could not appear, because the function that logs it was never called. The reps that read 1 were reading a phase chart armed minutes earlier by a batch that had not finished, which is why the gauge looked intermittent rather than absent.

Three things follow.

Arming is a property of the unit. baselineArmer (internal/controller/physical/unit_baseline_arming.go) is the one implementation and both brokers drive it. This adds no watchdog and moves no trigger: §Decision.4's mode interaction is unchanged, and the lease and the heartbeat still each fire their own. What changes is that both of them now load what they fire.

The repair is level-triggered, off what the runtime reports. The rule it replaces re-armed after a heartbeat failed and then recovered, which is a proxy for "the pod probably restarted" and misses everything that is not that: a dropped POST, a pod that restarts fast enough to answer the next beat, and the arm at loop start racing the runtime's own HTTP listener — measured twice in one day on the bench, refused the connection both times, with nothing retrying either. The runtime answers holdArmedProgram and holdBaselineArmed on the lease renewal and on the heartbeat, so the operator learns the edge posture on a transport that already runs at a sixth of the lease (a third when this was written, #1909) and repairs all of those under one rule. The question asked is about the baseline slot, not about whether anything is armed: a running phase's chart answers yes for the whole of that phase while the slot behind it stays empty, and the unit is uncovered from the instant the phase disarms.

A posture nothing can observe is a posture nobody maintains. Arming is process memory in the HoldController, so before this the only witness anywhere was dcs_runtime_hold_armed on the runtime's own hostNetwork metrics port. A drill asking "is anything armed at the edge" had to scrape a node address, and an operator asking it had no answer at all — which is why a defect this size sat in a shipped, bench-exercised feature for months. The control plane records the last report on Unit.status.runtimeBinding.edgeHold. Nil there means no acknowledged contact has produced a report, and a record with an empty program means the runtime says nothing is armed; those have opposite remedies and the status keeps them apart.

TestReconcileAvailabilityArmsTheBaselineInBothModes drives the whole decision from reconcileAvailability and fails on its Failover leg against the old shape. TestARestartedRuntimeIsRearmedByTheNextRenewal is the level-triggered half, with no failed exchange anywhere for an edge-triggered rule to notice, and TestAnArmedPhaseDoesNotStandInForTheBaseline holds the slot question apart from the "anything armed" one.

Amendment (2026-08-25, #1818): the latch is released by the command, not by the observation

§Decision.3 says the Autonomy watchdog fires "once per partition" and that recovery is an explicit ISA-88 Restart rather than a silent resume when the heartbeat returns. HeartbeatWatchdog.fired (internal/adapter/heartbeat.go) implements that, and it is correct. What was wrong is who cleared it.

Watchdog().Release() is the only thing anywhere that clears the latch. It is reached from POST /api/v1/hold/release, whose only caller in the product is PhaseReconciler.releaseEdgeHold, and that call used to require Phase.status.edgeSelfHeld. That field is set in exactly one place — reconcileEdgeSelfHold, whose evidence is an HTTP GET to the runtime pod. So the release was gated on the control plane having watched the edge hold happen, and on a partition it has not: the pod is on the far side of the fault.

The two clocks make that the ordinary case rather than a race. The control plane's own road to Held is reconcileUnitFaultHold, which fires at pod-not-ready plus a 30-second grace and sets no edge flag. The watchdog fires at holdGraceSeconds, 60 seconds by default, behind the partition. The loud road reaches Held first, every time. The Restart that follows found the gate false, sent no release, and left the latch set for the life of the runtime process.

Measured on the bench on 2026-08-25, two reps of the same partition drill on one runtime process. Rep 1 did everything this ADR promises: the watchdog fired at T0 plus 58.1 seconds, ran the unit's safeStateChart, and settled 0.3 seconds later. Rep 2 ran nothing — no watchdog line, no hold, no hold event on the edge's disk, no AuditRecord — and both outputs read HELD_LAST_VALUE for the whole partition while the unit still reported runtimeBinding.edgeHold.baselineArmed: true with program: unit-baseline and dcs_runtime_hold_armed still read 1.

Three things follow.

The command out of Held is the whole condition. releaseEdgeHold is sent on any ISA-88 command that leaves Held — Restart, Stop or Abort — and the edgeSelfHeld status fields are cleared only when this phase was carrying them. The alternative considered was to make the flag honest instead, by having reconcileUnitFaultHold poll the edge before reflecting. That is more faithful to what the field means and it was rejected as the fix, because it leaves the release depending on an observation that can still fail, and the failure is silent. The observation is worth having; it is not worth gating safety on.

A no-op release costs a request, and a missed one costs the next partition. handleHoldRelease clears the self-held state and the latch unconditionally, and a runtime that is not self-held has nothing to clear, so the request is cheap and idempotent. A command out of Held is not a hot path. This is the same trade §Decision.3 already makes in the other direction: fire once and make the operator ask for the resume.

Armed and able-to-fire are different facts, so they get different signals. Nothing in the product said the watchdog had spent itself. hold_armed, holdBaselineArmed and edgeHold.program are all about the PROGRAM, and all three answer the same in both states, because the chart really is armed and simply cannot be run. dcs_runtime_hold_watchdog_latched and Unit.status.runtimeBinding.edgeHold.watchdogLatched are the discriminator, reported on the heartbeat alongside the arming. The latch gauge is published only where a watchdog exists, so an absent series means "this unit has no edge watchdog" and 0 means "it has one and it can fire" — the #1648 rule, that a signal's absence must mean exactly one thing.

The status field obeys the same rule by carrying no omitempty. The gauge gets its absence-means-something from the mode, and a bool on a CRD has no such luxury: omitempty erases a false, and an erased false is byte-identical to what a control plane predating the field writes. A reader would have to settle which one it was looking at from the build, which is the confusion this field was added to end. So the value is always serialised, exactly as baselineArmed beside it is, and the schema stays +optional because an object stored before this amendment really does lack the key. TestEdgeHoldArmingWritesTheLatchWhenItIsFalse holds it, because re-adding omitempty compiles and passes everything else.

TestReconcile_Restart_ReleasesEdgeHold_AfterUnitFaultRoad drives Held by the unit-fault road and fails against the old gate. TestReconcile_NoReleaseWithoutACommandOutOfHeld is what stops it passing for the wrong reason: a release sent on every reconcile would satisfy the first test and would also clear a latch behind a partition that is still on, which is the silent resume this ADR forbids.

Amendment (2026-08-25, #1833): the arming reaches the surfaces a person opens

The amendment above records the last report on Unit.status.runtimeBinding.edgeHold, and the amendment before it adds the field that separates armed from able-to-fire. Neither put the answer anywhere a person could read it. grep -rl EdgeHold over the tree returned the adapter, the two controllers, the API types and their tests, and no file under internal/gateway/ or cmd/dcs/. So the question this whole ADR exists to answer — is this equipment covered right now — was reachable through kubectl get unit -o jsonpath and nothing else, four months after the record was written.

Placement was a founder call, taken on 2026-08-25 against measured mock-ups drawn in the shipped gateway CSS. Three surfaces carry it, and the split between them is which failure each one survives.

The Unit detail's Properties block carries an Edge Hold row, reading the control plane's record. It is the one that survives the partition the arming exists for, and it costs 42px of a 316px block. The row renders on every unit, healthy ones included, which is the opposite of the rule the three #1772 neighbours follow: those carry exceptions, so a permanent line saying nothing is wrong would be annunciation with no decision behind it (ADR 0029). This row exists so a commissioning engineer can confirm the promise holds, which a row that appears only when it is broken cannot do. Runtime Ready three rows above it is green on every healthy unit for the same reason.

dcs get runtime carries three lines, read live off the runtime pod's own HoldController and Watchdog through the diagnostics proxy. It is the freshest reading and the one that stops existing when the pod does. Where it disagrees with the row above, the disagreement is the finding.

dcs get units carries an EDGE HOLD column, off the same control-plane record. Neither of the other two closes the gap the issue opened with: one is a gateway page and the other reaches the pod, which on a partition is on the far side of the fault. A drill script and an engineer at a terminal have this.

The verdict is computed once, and one of its inputs is not on the report. Six readings live in edgeHold plus the unit's own spec.safeStateChart, and two of them are byte-identical on the wire. A report saying nothing is armed means a chart failed to arm on a unit that declares one, and means the configured posture on a unit that declares none — opposite remedies, and the second is not a fault at all. unitEdgeHoldToDTO is the one place that decides; the browser and the CLI render its summary and never re-derive it. dcs get runtime is the exception and has to be: the runtime holds no Unit spec, so its Armed Program line says what is armed and deliberately does not say whether that is wrong.

The latch outranks everything the report says about the program. On the #1818 bench rep the program was named and the baseline slot was armed while nothing either of them described could run. Ranking watchdogLatched below armed renders that unit green, which is the reading that lost a whole second partition. TestLatchedOutranksAnArmedProgram fails against a latch demoted below the program fields, and TestNothingArmedIsDecidedByTheDECLARATION fails against a verdict that stops reading spec.safeStateChart.

Amendment (2026-09-02, #1919): a hold is answered by the work that moves on, not only by a Restart

The #1804 amendment drew a line: a hold this runtime ran is released by an ISA-88 Restart and by nothing else, and only a hold it found on disk is retired on evidence of presence. The line is right about which evidence clears a live hold — a lease grant or a heartbeat proves the control plane is back, not that anyone has recovered the unit, and a hold sitting under a Held phase looks exactly the same the moment before its Restart as it would if nobody ever sent one. It was wrong about the premise underneath it, which is that every hold has a Restart coming.

Measured on the bench on 2026-09-02, drill 12 rep c of the #1894 verification, chart 0.7.1. The unit fenced for 9.5 s (#1909). The runtime ran its bounded hold, dcs_runtime_self_held went to 1 at 20:56:28.7, the lease was re-granted and the fence cleared at 20:56:38, and the gauge stayed at 1. The batch that was fenced had already completed before the cord came out, so the hold ran on an idle unit under the baseline. No phase was Held by it, so no Restart was ever going to arrive for it, and the next batch's Start is not a Restart. Two full batches then ran to Complete on that runtime, both regulating at setpoint, and the gauge still read 1 at 22:40. Rep 2's harness row carried self_held=min=1 max=1 through a rep in which nothing held, which is the cost: the signal could not have seen a second hold if one came.

What answers a hold. Two more roads, and both are the control plane's word on the same POST /api/v1/hold/release, which now carries the road in its body and records it on the released audit event:

  • later-phase-started — a phase starts on the unit and finds the runtime self-held under work that is not this phase's. A unit that takes a fresh Start has been recovered by that Start. The discriminator is edgeSelfHoldIsThisPhases, the reading the reflection already uses: a phase-scoped program names its phase, and a baseline hold that triggered before this phase started is not this phase's. A hold that is this phase's — the runtime fenced mid-Running and the ActionChart is re-entering after the outage — is left standing for the Restart or for the phase to end.
  • engaged-phase-ended — the engaged phase reaches Complete, Aborted, Stopped or Idle. Terminal is the other end of engaged, and there is nothing left to Restart. The release is unconditional on the transition, on the argument #1818 already made for the Restart's: one request on a transition that is not a hot path, and a no-op on a runtime that is not self-held.

The Start road reads GET /api/v1/status rather than riding the arm request, because the bench's phases declare no holdingChart and the arm request never leaves for a phase without one. A road that could not fire on the unit the defect was measured on would not have closed it. The read is one GET on a phase start; when it fails the hold is left for the next road rather than guessed at.

The watchdog latch shares the premise, and it is the safety half. #1818 made the Restart's release clear HeartbeatWatchdog.fired, on the same reasoning: recovery from an edge hold is a Restart. In Autonomy the shape measured here — a baseline hold on an idle unit, then a fresh batch — leaves the watchdog latched under a live batch, so the next partition runs no chart at all while every instrument reports the program armed. That is worse than the gauge. Both hold-side states are released together on every road; Failover has no watchdog and is unaffected on that half.

What is deliberately not done. The runtime does not clear itself on presence: RetireInheritedHold is unchanged and still refuses a hold this process ran, because presence cannot tell a hold awaiting its Restart from one nobody will Restart. The runtime refuses a release naming a road it does not know (400), so a misspelling in the control plane cannot write a road that does not exist into the audit trail. An absent body or an empty reason is the Restart, which is what every caller before this amendment meant. During a mixed-version rollout an older runtime ignores the body and records every road as control-plane-restart; an older operator's bodiless release lands as the Restart on a newer runtime. Neither leaves a hold standing.

The cost is the reflection's own: the Start road compares the runtime's triggeredAt with the phase's startTime, two clocks, and it abstains rather than refuses where the reading is unavailable. A hold triggered within clock skew of a phase's Start reads as that phase's own and waits for the terminal road, which is the fail-safe direction.

TestReconcile_Start_AnswersAHoldThatIsNotThisPhases drives the bench shape and fails against a tree with either road removed; TestReconcile_Start_LeavesThisPhasesOwnHoldForTheTerminalRoad is what keeps the Start road from answering a hold the phase is still under. TestHoldRelease_NamesTheRoadTheControlPlaneTook holds the runtime to recording the road and releasing the latch with it, and TestHoldRelease_RefusesARoadThatDoesNotExist holds the refusal.

Amendment (2026-09-03, #1928): the audit bridge is subscribed by the leader, and every replica connects under its own id

Context. The chart has run the physical-operator at two replicas since #1893. Both replicas built their MQTT client at process start under the literal id physical-operator, and both subscribed the hold-lifecycle and safe-stop bridges before leadership was known. MQTT requires a client id to be unique per session, and Mosquitto enforces it by closing the older session when a CONNECT arrives under its id. The bench broker logged 495 New client connected lines in 70 minutes, alternating between the two pod IPs, with already connected, closing old connection before each. The leader's subscription to the hold-event topic was dropped and remade at that cadence, and a hold event published into one of those gaps reached nobody.

Decision.

  • The client id is the pod's. Every operator names its session after POD_NAME (the downward API, set by the chart) and falls back to the hostname, which the kubelet sets to the pod name. The bare component name never reaches the broker: with neither, the id carries a random suffix.
  • Every replica connects at start. The connection is not the leader's, because a replica taking over needs a session already up: its first unit-state announcement must not wait on a TCP, TLS and CONNECT round trip inside the grantor gap (#1909). The standby's connection is idle.
  • The subscriptions are the leader's. The two bridges are subscribed by a leader-election Runnable (AuditBridgeSubscriber) on winning the election and unsubscribed when the manager stops. A subscription is a writer: each event becomes an AuditRecord under a generated name, and two replicas holding it would write every hold and every safe stop twice. Before this amendment the trail stayed single only by the accident this amendment removes, because whichever replica held the connection at that instant was the one materialising. A standing-down leader (#1894) keeps running after its manager stops, and the unsubscribe on stop is what keeps it from writing beside its replacement.

What this does not change. An event published while no leader holds the subscription is received by nobody and kept by nothing, and the election gap is one instance of that. The runtime's queue (#1779) survives a lost broker on the publisher's side; it does nothing for a subscriber that was not there. That is #1927, and a stable session the broker could resume for the leader role is not its answer here, because a stable id is exactly what two live processes cannot share.

Consequences. The broker log shows one New client connected per pod start and no already connected lines. dcs_mqtt_connected{component= "physical-operator"} is read per pod through Prometheus's own pod label; the standby's reading is its idle connection, which is true. The chart gate liveness-grantor.sh asserts POD_NAME comes from the downward API and not from a value in the template, and TestMQTTClientConfig_TwoReplicasCarryTwoClientIDs holds the binary to it.

Amendment (2026-09-03, #1927): the durable session the leader alone holds is the reader that outlasts the absence

The #1779 amendment made the runtime's half of the edge audit path a disk write, so a hold never waits on the network to reach its safe state. The runtime's half held. The record still did not survive.

Measured on the bench on 2026-09-03, chart 0.7.3, the first grantor-absent rep. The physical-operator was scaled to 0 for 150 s with an idle Failover unit. The runtime did what ADR 0006 asks — bounded edge-local hold, settled into safe state, self-fenced, fenced for 132 s — and published every one of those events to a broker that was up throughout. The trail records none of them. It records only the released, 37 s after the operator was back, so it says a hold was released that it never says ran.

The operator is the only thing that materializes these events into AuditRecords (the runtime never writes the apiserver). It had subscribed with a session that ended when its connection did, so the broker held nothing for it between connections, and the events published while it was absent went to a broker with no subscriber and were gone. This is not the #1780 store-and-forward gap: the broker was up, the runtime published, its queue had nothing to hold. The loss is on the subscribing side, and it falls on exactly the fence that matters most — the one shape that fences a Failover unit is the grantor going away (#1893), and every fence of that class runs while the only reader of its record is the absent operator.

The fix is a durable subscriber. The two audit bridges — the hold lifecycle and the terminal-stop safe state (#1283) — move onto their own MQTT connection that presents a stable role client id (physical-operator-audit-bridge) with a week-long SessionExpiryInterval and clean-start false. The broker keeps the session, and every QoS 1 message published to it while the operator is away, until the next leader reconnects under the same id and drains it. Three things make that connection what it is, and each is load-bearing:

  • It is the leader's alone (NeedLeaderElection). A session is keyed on the client id, so two live connections under it are a takeover, not two subscribers, and the standby reconciles nothing to give a record to.
  • Its id is a role, not a process. The pod that comes back after an eviction has a new name, and a session keyed on the pod name would be orphaned by the very restart it exists to survive. This is the opposite of what the publisher on the same operator needs, where an id per process is what stops two replicas taking each other over (#1928) — which is why the audit bridge is a second connection and not a setting on the first.
  • Its handlers are registered before the connection starts. A resumed session's queued messages arrive on the heels of the CONNACK, and paho drops an inbound publish that no registered handler matches, so a handler registered after Connect returned could miss the replay the session was kept for.

The broker's own max_queued_messages bounds how many messages it will hold for that session; the chart raises it from Mosquitto's default of 1000, since a grantor loss fences every Failover unit at once and a hold is a dozen records each.

The cost is at-least-once delivery on a resumed session: a message this pod acknowledged to the apiserver but not to the broker when it died is redelivered to its successor and becomes a second AuditRecord. The recorder already takes that trade on its own retries, where a duplicate is "strictly preferable to a drop," and it is the same trade here.

Between the runtime's store-and-forward queue (the broker being away) and this durable session (the subscriber being away), the record now has a home at every point on the path.

TestIntegration_ASessionOutlivesItsSubscriber drives the bench shape against the pinned broker. A subscriber leaves, a publish lands while it is gone, and a new process under the same id comes back. It runs both arms, so the durable session's delivery is proven against a plain session's loss. TestAuditBridgeSubscriber_SubscribesBeforeItConnectsAndDisconnectsOnStop holds the register-before-connect order and the disconnect-not-unsubscribe on losing leadership, and NeedLeaderElection keeps it the leader's alone. TestAuditBridgeClientConfig_IsADurableSubscriber holds the session and the role id. TestAutopahoConfig_CarriesTheSession holds that the session fields reach the wire and that a client asking for no session is handed none. The annunciation half, that the same absence raised no alarm for the fence, belongs to #1890 and is recorded there.

Amendment (2026-09-11, #2137): a report belongs to the runtime that made it, and the record follows the report

The #1806 amendment put the runtime's last arming report on Unit.status.runtimeBinding.edgeHold, and the #1833 amendment put that record in front of a person. Two things about the record were left to the reconcile loop, and a manual failover of an Autonomy unit found both.

The first is whose report it is. The heartbeat loop keeps one memory per unit: what the runtime last answered, and which chart this loop landed on it. A re-bind moves the loop onto a different pod at a different address and the memory came with it. performFailover rebuilds the binding with no record, and the next pass copied the loop's memory of the previous runtime into it, stamped now, as the promoted runtime's own answer. On the multi-node capture stack that read armed for the 21 s between the re-bind and the broker's arm landing, over a runtime that had nothing armed. On the bench the previous runtime's last decoded answer was the one taken before its own arm landed, so the same copy read nothing armed, and the loop's memory that it had already armed that chart kept the optimistic arm from being posted to the new pod. A runtime's arming is process memory. Nothing learned from one process holds for the next, so a re-target onto a different address drops the report, the arming memory and the last-acknowledged instant together, in both brokers. The in-place road, a pod that came back at another address under the same binding, clears the record on the Unit for the same reason it already clears leaseEstablishedBy.

The second is when the record moves. It is written only when the reconciler runs, and an Idle unit has no requeue of its own, so a beat that learned the promoted runtime was armed waited on whatever next happened to reconcile the unit. Both brokers now nudge the reconciler when the answer moves, on the same channel a lease transition already uses. And the answer moves sooner: a re-target and the pod's readiness edge both arm and beat at once rather than at the next tick, which is a third of the grace away, and a beat that lands an arm asks once more so the record follows the act by seconds. The window an operator saw nothing armed over a standby that was armed is the tick plus the next reconcile; it is now the round trip.

TestAManualFailoverRearmsTheBaselineOnThePromotedRuntime drives the shipped road through the re-bind on both transports the channel has, HTTP/1.1 and HTTP/2 over TLS with the operator's own transport shape, and holds the three claims apart: the record says nothing until the promoted runtime answers, the arm lands inside one beat interval of the readiness edge, and the reconciler is nudged when the report moves. Each was proved by breaking what it guards.

Amendment (2026-09-22, #2155): a marker the phase has already answered is the hold being recovered from

The #1723 road to Held reads Unit.status.conditions[HeldByRuntimeFault] and reflects it onto the phase running on the unit, and the #1818 amendment settled that the command out of Held is what answers a hold. The two met on the bench on 2026-09-22 at 17:51:40Z, on the cord-pull rep of the manual-failover take. A confirmed-fenced failover of bench-loop-01 landed at epoch 122, the runtime came Ready on the standby, and the drill issued the documented Restart to the batch ten seconds later. Within one second the physical operator took the Restart (Held → Restarting → Running, marker removed), the procedural operator took the phase's (Held → Restarting), and then reflected the unit-fault hold onto the phase twice, and the tree held with it (#1823). The batch sat Held and the drill aborted it after 90 s. The same recovery had succeeded on every reboot-road pass that hour.

The physical layer removes the marker when the Unit leaves Held, in the same status write as the state. The phase controller is a different process reading the Unit through its own informer, and its post-Restart pass ran before that write reached its cache. So it read the marker the Restart had just answered as a fresh fault. The race is by construction, and whether it is lost is watch latency.

Three things follow.

The reflection records which marker it answered. Phase.status.reflectedUnitHoldAt is the marker's own LastTransitionTime, written by reconcileUnitFaultHold when it drives Held. runtimeFaultHold, the one lookup both the reconciler and the SFC status publisher read, refuses a marker stamped at or before it. A new fault always arrives with a later stamp, because the physical layer removes the condition when the Unit leaves Held and a condition that was absent is stamped afresh when it is set again. So the record cannot mask one.

The anchor is the marker reflected, not the moment the phase left Held. The issue proposed the latter, and it was rejected because of a road it breaks. A phase Held for another reason (a sync barrier, a coordination wait, an operator Hold) while the fault arrived never reflected the marker. By that anchor its resume would refuse the marker as older than the Restart and run the phase over held equipment. TestPhaseReconcile_RestartOutOfAnotherHold_StillReflectsTheFault is that road. The record is cleared by Reset, because the next Start is a fresh run against whatever the equipment says then.

The UnitProcedure's barrier reads a stamp of its own, the other way round. convergeUnitState leaves a Unit the physical layer holds where it is, awaiting the explicit Restart. That Restart is the UnitProcedure's own command out of Held, now recorded in status.leftHeldAt. A Unit still Held under a marker stamped strictly before it is a Unit the Restart never reached, which is the lost command the #1151 repair exists for, and the repair sends it again. The comparison is strict because both stamps carry second precision, and the safe reading of a same-second marker is a new fault. This read was already live rather than cached (#1774), so the stale-cache half of the defect never applied to it. What applied was a barrier that read every marker as awaiting a Restart that had already been given.