ADR 0008: Edge-local holding logic — partition-triggered safe-state sequencing at the unit runtime¶
Status: Accepted Date: 2026-06-16 Issue: #578 Related: ADR 0006 (refines its partition-gap context), ADR 0007 (extends its three-layer protection model), #566 (edge redundancy scoping)
Context¶
ADR 0007 established a three-layer protection model and was explicit about which layer survives a control-plane partition:
| Layer | Where it runs | Survives partition | Role |
|---|---|---|---|
Device interlock (DO/AO ILCK) |
FB scan, edge | Yes | Force one output to one static safe value while a trip is true |
Phase SFC guard (interlock: true) |
Procedural operator, control plane | No | Procedural safe-state sequencing — ordered multi-output response (vent first, then cut heat), parking the procedure in a resumable held step |
AlarmDefinition.exceptionAction |
Alarm controller, control plane | No | Annunciation, audit, batch-level Hold/Stop/Abort |
The gap #578 names is the middle row. A device interlock forces a single
output to a single value — it cannot sequence. The only layer that can run a
deliberate, ordered safe-state response (the thing a process actually needs:
close the charge valve, then stop the agitator, then open the vent) is the
phase SFC, and that executes in the procedural operator on the control plane
(internal/controller/procedural/phase_controller.go instantiates the engine;
each scan's READ/WRITE builtins cross HTTP to the runtime via
pkg/stbridge). On a control-plane partition:
- the edge FB scan keeps regulating to whatever setpoints were last commanded (the autonomy claim — accurate, ADR 0006), but
- the phase SFC stalls mid-step at an arbitrary point, so the equipment is frozen wherever the sequence happened to be — mid-charge with the charge valve held open is frozen, not safe, and
- no sequenced holding response can execute at the edge, because that logic lives in the control plane.
ADR 0006 made this precise and accepted it as the v1 posture: "Field outputs
hold last value (or device fail-safe) during the gap, exactly as for a pod
restart today" (ADR 0006 §Decision.3). The compliance traceability says the
same — docs/compliance/isa88.md Clause 7.4 row: "Held outputs hold last
value / device fail-safe during the gap." Frozen, not driven to a deliberate
safe state.
The pod-restart half of that sentence was not true until #1776. A restarting runtime rebuilt every block empty and wrote a compile-time default on its first scan, so a hold this ADR had armed correctly was discarded by the same fault's own restart. See ADR 0080.
Competitive + standards baseline (researched 2026-06-12, #578):
- DeltaV's shipped default on Batch Executive ↔ controller comm loss is the phase placed in Held via a watchdog, with Holding logic executing locally in the controller with full I/O access — deliberate safe-state actions, not frozen outputs (Emerson Batch Executive PDS). Rockwell FactoryTalk Batch documents the same watchdog → HELD philosophy. Note this is the non-redundant comm-loss path; it is orthogonal to controller redundancy (ADR 0006), which is a separate feature.
- ISA-88 Part 1 deliberately does not mandate execution placement (Clause 6.6.3: recipe/equipment separation is logical, physical separation optional; Clause 6.6.4: equipment procedural element internals are out of scope). What it expects is the Clause 7.4 exception response: malfunction → Hold to a known safe state → Restart with recipe-defined restarting logic.
So the gap is not "move the phase engine to the edge" — that would put mid-phase durable state on the node ADR 0006 deliberately allows to die, and complicate the 21 CFR Part 11 audit trail. The gap is narrow and specific: the edge has no locally executable, sequenced holding response.
Five constraints from the existing code shape the option space:
- One writer owns the drivers. ADR 0007 rejected a separate watchdog goroutine forcing driver writes precisely because "two writers to one output address race each other; the FB scan already owns output ordering and the one-writer guarantee." Any edge holding response must respect this — it cannot become a second driver writer.
- SFC encapsulation rule. All phase state logic must be SFC charts; no bare flat ST (project convention). A sequenced hold response is phase state logic.
- The SFC engine is already standalone.
pkg/sfcimports only the CRD types, the ST interpreter, and a clock; it accepts an injectableStepHandlerand supportsWithRestoredState. It does not today have a driver-bound execution path at the edge — itsREAD/WRITEresolve over HTTP in the control plane. Failoverself-fence and a local hold would fight. ADR 0006'sFailovermode self-fences on lease expiry — it stops FB output writes so a standby can take over (internal/adapter/lease.gofencedDriverdrops writes). A holding response that writes during that window contradicts the fence and risks dual writers against the standby.- Durable Held state lives in etcd, and the phase self-holds anyway. Phase
state lives in
Phase.Status(ADR 0006); the procedural operator's existing 60 s grace → self-Hold path (phase_controller.go) independently drives the phase to Held when it sees the runtime unreachable. Whatever the edge does must converge with that, not diverge from it.
Decision¶
The unit runtime gains an armed local-hold program: an SFC chart executed by an SFC engine embedded in the runtime when the edge's control-plane heartbeat watchdog fires. The hold chart drives a deliberate, sequenced safe state by writing the unit's own control-module tag space — the FB scan remains the sole I/O writer, so no new write path and no dual-writer is introduced. This makes ADR 0007's procedural safe-state layer partition-tolerant. Durable Held state continues to live in the control plane (ADR 0006); recovery is a normal ISA-88 Restart (Clause 7.4) — no silent resume. This refines ADR 0006's partition gap; it does not supersede it.
The scoped mechanism:
- What is armed. Two sources, in priority order:
- Per-phase (preferred): at phase start the procedural operator
downloads the active phase's
HoldingChartto the runtime as the currently-armed local-hold program — the same chart that runs the ISA-88 Holding state in the control plane, now also staged at the edge. This reproduces the DeltaV default (phase-specific holding logic) and is correct precisely because a mid-charge partition arms the charge phase's holding logic. - Unit baseline (fallback): a new
UnitSpec.safeStateChart(SFC), deployed with the unit's control program and always armed when no phase is active. This closes ADR 0007's idle-unit / manual-mode coverage gap for the sequenced case, the same way device interlocks closed it for the single-output case.
Edge-armable constraint. A chart is armable at the edge only if it
READ/WRITEs the unit's own control-module tag space (no cross-unit, no
control-plane-only data). A chart referencing out-of-edge-scope data cannot
be staged; the runtime falls back to the unit baseline and the operator
surfaces a warning. A lint/validation enforces this.
-
Execution model. The runtime embeds the
pkg/sfcengine with aStepHandler/tag binding whoseREAD/WRITEresolve directly against the local control-module tag space (the FB network's variable bindings) rather than over HTTP. The hold chart manipulates FB inputs (setpoints, valve commands) exactly as the phase SFC does today viapkg/stbridge; the FB scan regulates to them and remains the only writer to drivers. Device interlocks (ADR 0007) still sit underneath and can override even the hold chart's commanded outputs. The SFC encapsulation rule is honored — the hold response is an SFC chart. -
Trigger contract — a unified control-plane heartbeat. The local-hold trigger is "no control-plane heartbeat within the timeout," evaluated by a purely local timer at the edge (no network reach required to decide):
- In
Failovermode the heartbeat is the existing lease renew (/api/v1/lease/renew,internal/adapter/lease.go) — reuse it, do not add a second watchdog. - In
Autonomymode there is no edge watchdog today (the lease guard is inert). Add a lightweight liveness heartbeat the procedural/physical operator issues each reconcile (carrying "phase X is live") — no fencing, no lease semantics, just a timestamp the edge watchdog consumes. -
Timeout is the unit's
availability.leaseDurationSecondsinFailover, and a newavailability.holdGraceSeconds(default mirrors the control plane's 60 s phase self-hold grace) inAutonomy. -
Mode interaction (resolves constraint 4).
Autonomy: on heartbeat loss past the grace window, the edge runs the armed hold chart to drive a deliberate safe state, then keeps holding (the FB scan continues regulating to the safe-state setpoints the chart established). No fence — the edge is the legitimate single writer.-
Failover: on lease expiry the edge runs the armed hold chart as a bounded safe-state sequence, then self-fences (stops writes) so a standby can take over cleanly. The standby re-bind already happens only after lease-expiry + margin (ADR 0006); the margin must be ≥ the hold sequence's bounded duration, so the partitioned node finishes its safe-state actions and fences before the standby writes — no dual writer. A node that is truly dead (not merely partitioned) cannot run the hold; that case falls back to ADR 0006's existing behavior: outputs at device fail-safe / last value until the standby re-binds and Restart drives them. What the re-bind itself writes is ADR 0082, and until then it was a compile-time default rather than a continuation. Local-hold strictly improves the alive-but-partitioned case; it does not regress the dead-node case. -
Reconciliation on reconnect. The runtime reports, via its status (MQTT status + a gRPC/HTTP field), that it self-held, which program it ran, and when. On reconnect the operator reads this and reflects Held on the Phase/Unit state machines. This converges with the control plane's own 60 s self-hold path (constraint 5): during a partition both sides independently move toward Held — etcd marks the phase Held, the edge holds outputs — and on reconnect they agree. Recovery requires an explicit ISA-88 Restart (Clause 7.4); there is no silent resume.
-
Audit at the edge. Local-hold executes while the control-plane audit path is unreachable. The runtime emits hold-lifecycle events (armed, triggered, each safe-state action, self-fenced) onto the existing MQTT store-and-forward queue (
internal/adapter/queue,/var/lib/dcs/runtime/queue/), replayed on reconnect and materialized intoAuditRecordCRs by the audit bridge. This preserves the "audit must not block the control loop" rule and reuses the historian store-and-forward precedent rather than giving the runtime direct apiserver write access.
Onto the queue always, never around it (#1779). The first
implementation of this sink tried a live publish and fell back to the queue,
which is the right shape for telemetry and the wrong one here. Hold events
are emitted on the goroutine driving the equipment to its safe state: the
triggered event goes out ahead of the engine, and each safe-state action
goes out from inside the step that performed it, which does not advance
until its own emission returns. A live publish at QoS 1 waits for a PUBACK,
and during a control-plane partition the broker's answer is exactly what
nothing can promise. Measured on the bench, a hold that settles in 300ms
left the field outputs at their pre-fault values for 10.3 seconds. The sink
therefore appends to the queue and wakes the replay worker, which does the
network part on its own goroutine. Bounding the publish was rejected as the
fix, because a bound still spends the bound. The queue is also the more
honest record: the event is durable before control moves on, where a live
publish still in flight when the pod dies leaves nothing behind.
The bench rep ran in Autonomy, where the hold has no deadline and the
delay costs 10.3 seconds of equipment sitting at its pre-fault setpoint. In
Failover it costs more than that. The hold there is bounded at half the
lease, clamped to between 5 and 30 seconds, and the stall sits ahead of the
engine rather than inside it. A stall long enough to consume that budget
would have the runtime self-fence having written nothing at all, and a
standby would take over the frozen unit that the bounded hold exists to
prevent.
The trail carries the order, and the timestamp never could (#1813).
This decision is about a sequenced safe state, so the order of the
actions is the content of the record and not a detail of it. The recorded
timestamp cannot carry it. AuditRecordSpec.timestamp is a metav1.Time,
which serialises at whole-second precision, while the hold engine scans
every 100 ms, so the two WRITEs of a two-step safe state land on one
second and tie. The gateway then sorted the trail on that value with an
unstable sort, and two reads of one unchanged trail could disagree about
which output was driven to its safe value first. Each event therefore
carries an emission ordinal, assigned by the runtime under its own lock at
the moment of emission rather than derived from a clock. That distinction
matters here more than anywhere else in the product. The node is
partitioned, its clock may step when the network returns, and the ordinal
is unaffected by either. The ordinal rides the hold snapshot, so a pod that
restarts mid-partition resumes numbering where it left off instead of
renumbering over the events already queued. The published guarantee and the
ordering every audit surface applies are recorded in
21 CFR Part 11 Compliance Traceability.
Alternatives Considered¶
- Download a compiled safe-state FB network instead of an SFC chart (engine already at the edge — issue option 2). Rejected as the primary: a flat FB network cannot sequence (vent-then-cut-heat), which is the entire gap — ADR 0007 keeps sequencing in the SFC layer for exactly this reason, and the FB layer already provides the static single-output safe state via device interlocks. Re-encoding ordered sequences as FB logic would duplicate the SFC engine's job and violate the SFC encapsulation rule.
- Move the whole phase/SFC engine to the edge. Rejected: it puts durable mid-phase state (active steps, fired transitions, variable values) on the node ADR 0006 deliberately allows to die, breaks the etcd-resident phase state that lets a running phase survive failover, and complicates the Part 11 audit trail. The fix is a fallback safe-state responder at the edge, not relocating sequencing ownership.
- Unit-level static safe-state chart only (no per-phase arming). Rejected as the sole answer: it loses phase-specific safe states (mid-charge wants a different response than mid-heat), which is the DeltaV parity the issue is chasing. Kept as the baseline for the idle/no-phase case.
- Status quo — frozen outputs + device interlocks only (ADR 0006 v1 posture). Rejected as terminal: a device interlock cannot sequence, and frozen-at-last-value is not a deliberate safe state. The diligence comparison against the DeltaV/Rockwell shipped default fails on exactly this point.
- A separate watchdog goroutine forcing driver writes at the edge. Rejected for the same reason ADR 0007 rejected it for interlocks: two writers to one address race; the FB scan owns the one-writer guarantee. The embedded SFC engine writes the FB variable space, never the drivers directly.
Consequences¶
- API / code surfaces that move:
UnitSpecgainssafeStateChart(SFC, the unit baseline) andavailability.holdGraceSeconds(Autonomy trigger timeout).- The unit runtime embeds the
pkg/sfcengine with a driver-bound tag handler (internal/adapter/,cmd/unit-runtime/); a new edge watchdog consumes the heartbeat; new self-held status field on the adapter API. - The adapter API gains an arm-hold-program endpoint (download the active
phase's
HoldingChart) and, forAutonomy, a liveness heartbeat endpoint;Failoverreuses/api/v1/lease/renew. - The procedural operator downloads the phase
HoldingChartat phase start and reconciles edge-self-held →Heldon reconnect (phase_controller.go). - The audit bridge consumes edge hold-lifecycle events from the MQTT queue
and materializes
AuditRecordCRs. - New metrics: hold armed/triggered counters, hold-sequence duration, edge-self-held gauge.
- Compliance:
docs/compliance/isa88.mdClause 7.4 row (control equipment malfunction → Hold → Restart) is refined: the partition response becomes a deliberate sequenced safe state, not frozen-at-last-value. The "Held outputs hold last value" caveat narrows to the dead-node / no-armed-chart case.docs/library/alarms-and-interlocks.md's three-layer table gains a note that the procedural safe-state layer is now partition-tolerant at the edge.- This is a candidate execution vehicle for the existing
ProcessExceptionsafeState.structuredTextrow (currently "Partial — validated and logged, not executed against equipment"): an armed unit baseline chart is where a process-exception safe state could actually run. Noted for the implementation epic, not decided here. - 21 CFR Part 11: the edge-buffered, replayed audit trail records the hold actions taken while the control plane was unreachable.
- Docs that update when this ships:
docs/ha-failure-modes.md(partition row: deliberate safe-state, not frozen),docs/architecture.md(autonomy section),docs/adr/0007three-layer narrative cross-link. - Marketing posture: now that this has shipped (epic #597, merged), the accurate partition claim is "watchdog-triggered local holding logic drives a deliberate safe state at the edge during a control-plane partition" — the DeltaV/Rockwell default reproduced. The pre-ship claim ("autonomy + restart-replay + supervised Hold, outputs frozen at last value") now applies only to the dead-node / no-armed-chart case. Never claim bumpless redundancy (that remains explicitly rejected, ADR 0006).
- Default behavior unchanged: with no
safeStateChartand no phaseHoldingChart, the edge behaves exactly as today (frozen outputs). The feature is additive and opt-in per unit/phase. - Reversibility: moderate now that this has shipped —
UnitSpec.safeStateChartis a public CRD contract and the heartbeat is a runtime API surface. (It was high before implementation, when this was still a design posture.) - Follow-ups (implementation sub-issues): embed driver-bound SFC engine at
the edge;
UnitSpec.safeStateChart+holdGraceSecondsAPI; Autonomy liveness heartbeat + edge watchdog; per-phaseHoldingChartarming at phase start;Failoverhold-then-fence ordering + margin guard; edge-self-held status + reconnect reconciliation toHeld; edge audit buffering of hold events; edge-armable chart validation/lint; docs + compliance updates.
Amendment (2026-08-10, #1406): edge-armable covers capabilities, not only data¶
The edge-armable constraint in §Decision.1 was written about data — the
chart may address only the unit's own control-module tag space. It said
nothing about capabilities, and the two fail the same way. Eleven of the ST
dialect's builtins cannot execute at the edge at all, because each one needs
something the partition has taken away: an operator (PROMPT,
PROMPT_CHOICE, PROMPT_VALUE), the gateway an external system delivers a
measurement through (AWAIT_RESULT,
ADR 0055),
the apiserver (MODE), the phase state
machine the control plane drives (COMMAND), control-plane state about where
the action chart stopped (STEP_ACTIVE), or a network session to another
machine (CALL_SERVICE and the three MTP builtins, ADR 0045).
A chart calling one of them passed every gate, ran correctly on the control-plane path every time it was exercised, and errored mid-sequence only during the partition it was armed for. That is the worst moment to discover it, and it was discovered as an ST runtime error inside a hold, with no operator and no control plane to report it to.
The decision is unchanged and its scope is stated fully. Such a chart is not armable, which is the outcome §Decision.1 already defines: the runtime refuses to stage it, the phase falls back to the unit baseline safe-state chart, and the operator is told — while the control plane is still up to be told through. Authoring one remains legitimate, because a holding chart that prompts the operator, or commands a PEA service to hold, is exactly right for the ordinary control-plane hold it will normally run in. What changes is that its edge cover is one posture coarser, and that fact now arrives at arming time rather than during the partition.
ValidateEdgeArmableBuiltins (api/procedural/v1alpha1/edge_armable.go) is
the static check, the runtime applies it when a chart is staged
(internal/adapter/hold.go), and make lint-edge-armable enforces it over the
example corpus alongside the tag-space half. Because the two halves must not
drift, TestEdgeUnavailableBuiltinsMatchHoldRuntime runs every builtin the ST
package implements through the real hold environment and fails if the list and
the runtime disagree in either direction.
Amendment (2026-08-24, #1804): a recovered snapshot is inherited, and what it says is about a program¶
§Decision.5 says recovery requires an explicit ISA-88 Restart and that there is no silent resume. That is right about a hold this runtime ran. It was applied to a hold this runtime only found, and those are not the same claim.
HoldController.Recover restores the self-held flag from disk on startup, so a
pod that restarted mid-partition reports the truth the moment it comes up. What
it cannot know is whether the partition that raised the hold is still on. If it
is, the watchdog or the lease re-triggers RunHold and this process owns the
hold. If it is not — the pod came up into a working control plane, carrying a
snapshot of a hold that ended — then the road §Decision.5 names has no traveller:
POST /api/v1/hold/release rides a Restart, and the batch that produced the
snapshot is finished. There is nothing to restart.
Measured on the bench during the drill-9 sitting for #1793. bench-loop-01's
runtime pod came up at 17:22:49 carrying drill 3's 16:51 self-fence, three
unrelated batches then ran to Complete on it, and dcs_runtime_self_held read
1 through all three — 361 scrapes at 1 Hz, no failures — while the loop
regulated at 59.96–60.06 % PV and dcs_runtime_fenced stayed 0.
Two things follow, and they are separate defects with a shared root.
The runtime retires what it inherited. The flag now records whether it was
raised here or found on disk. A hold this process ran is released by a Restart
and by nothing else, unchanged. A hold it inherited is retired by
RetireInheritedHold on evidence that the control plane is running this runtime
again: a lease grant in Failover, a heartbeat in Autonomy. Either one says the
road a Restart travels is open, so a claim still standing on it is about work
that is over. The released audit event names which road was taken —
control-plane-restart or inherited-snapshot-superseded — so the trail can
tell an operator's recovery from a retirement.
The cost is stated rather than argued away. A pod that restarts mid-partition,
whose partition then heals in the seconds before the control plane polls
GET /api/v1/status, loses the reflection. It does not lose the Hold, which the
control plane's own 60 s self-hold path lands on a phase whose runtime was out of
reach that long (§Decision.5's convergence, from the other side), and it does not
lose the record, which is §Decision.6's hold-event trail. What shipped instead
was a signal pinned at 1 for the life of a pod, which fails towards alarm rather
than towards all-clear and is useless in both directions — #1646's rule in its
quieter shape.
A reflection is about a program, not about a flag. reconcileEdgeSelfHold
drives the phase to Held without running its HoldingChart, on the reasoning
that the edge already ran it. That reasoning belongs to the hold's own phase.
Taken at face value, a self-held flag left over from other work drove a healthy
running phase to Held having established no safe state at all, told the operator
its runtime self-held during a partition that was not happening, and named a
phase from a batch half an hour gone. The consequence is not confined to
Autonomy: shouldPollEdgeHold fires on any phase that is Held or has lost
contact with its runtime, in either availability mode.
edgeSelfHoldIsThisPhases is the discriminator, and it takes two readings
because the program name settles only one case. A phase-scoped program names its
phase outright and phase names are unique per batch. The unit baseline names
nothing — a partition during a phase with no HoldingChart legitimately runs it —
so time settles that one: a hold triggered before this phase started running is
not this phase's hold. Both readings abstain rather than refuse where they cannot
tell, which is §Decision.5's direction: reflecting a hold that turns out not to
be ours costs a Restart, and missing one costs the reflection this ADR exists for.
TestReconcile_EdgeSelfHeld_IgnoresAnotherPhasesHold drives the bench snapshot
in front of a later batch's phase and fails against the old shape;
TestReconcile_EdgeSelfHeld_ReflectsABaselineHoldFromThisRun is what keeps the
guard from being a blanket refusal. On the runtime side,
TestRetireInheritedHold_LeavesAHoldThisProcessRan and
TestRetireInheritedHold_RunHoldReclaimsARecoveredSnapshot hold the line the
retirement road must not cross.
Amendment (2026-08-24, #1806): the baseline is armed by the unit, not by the transport that carries liveness¶
§Decision.1 says the unit baseline is "always armed when no phase is active",
and §Decision.4 says separately that a Failover unit runs "the armed hold
chart" as a bounded sequence on lease expiry. Both are about the unit. The
implementation put the arming inside the Autonomy liveness heartbeat loop —
§Decision.3's transport — and reconcileAvailability stops that loop in
Failover mode, so UnitSpec.safeStateChart reached a Failover runtime on no
road at all. The two paragraphs above were describing a posture the product had
for one availability mode out of two.
What the gap costs depends on the phase. A running phase that declares a
holdingChart arms it in preference and covers the window by accident, for as
long as it runs. A phase that declares none — which is the ordinary case, and
the case the bench drill for the unit chart exists to measure — had nothing
armed, and the bounded pre-fence hold §Decision.4 promises ran a program that
did not exist. The posture was frozen outputs, which is precisely what §Context
rejects as terminal.
Measured on the bench, whose one unit is Failover with a safeStateChart
declared. dcs_runtime_hold_armed read 0 at every drill rep that took the
reading against a runtime no phase had armed, and the diagnosis the harness
offers an operator — check the physical-operator log for arming unit baseline
hold chart failed — could not appear, because the function that logs it was
never called. The reps that read 1 were reading a phase chart armed minutes
earlier by a batch that had not finished, which is why the gauge looked
intermittent rather than absent.
Three things follow.
Arming is a property of the unit. baselineArmer
(internal/controller/physical/unit_baseline_arming.go) is the one
implementation and both brokers drive it. This adds no watchdog and moves no
trigger: §Decision.4's mode interaction is unchanged, and the lease and the
heartbeat still each fire their own. What changes is that both of them now load
what they fire.
The repair is level-triggered, off what the runtime reports. The rule it
replaces re-armed after a heartbeat failed and then recovered, which is a proxy
for "the pod probably restarted" and misses everything that is not that: a
dropped POST, a pod that restarts fast enough to answer the next beat, and the
arm at loop start racing the runtime's own HTTP listener — measured twice in one
day on the bench, refused the connection both times, with nothing retrying
either. The runtime answers holdArmedProgram and holdBaselineArmed on the
lease renewal and on the heartbeat, so the operator learns the edge posture on a
transport that already runs at a sixth of the lease (a third when this was
written, #1909) and repairs all of those under one rule. The question asked is about the baseline slot, not about
whether anything is armed: a running phase's chart answers yes for the whole of
that phase while the slot behind it stays empty, and the unit is uncovered from
the instant the phase disarms.
A posture nothing can observe is a posture nobody maintains. Arming is
process memory in the HoldController, so before this the only witness anywhere
was dcs_runtime_hold_armed on the runtime's own hostNetwork metrics port. A
drill asking "is anything armed at the edge" had to scrape a node address, and
an operator asking it had no answer at all — which is why a defect this size sat
in a shipped, bench-exercised feature for months. The control plane records the
last report on Unit.status.runtimeBinding.edgeHold. Nil there means no
acknowledged contact has produced a report, and a record with an empty program
means the runtime says nothing is armed; those have opposite remedies and the
status keeps them apart.
TestReconcileAvailabilityArmsTheBaselineInBothModes drives the whole decision
from reconcileAvailability and fails on its Failover leg against the old
shape. TestARestartedRuntimeIsRearmedByTheNextRenewal is the level-triggered
half, with no failed exchange anywhere for an edge-triggered rule to notice, and
TestAnArmedPhaseDoesNotStandInForTheBaseline holds the slot question apart
from the "anything armed" one.
Amendment (2026-08-25, #1818): the latch is released by the command, not by the observation¶
§Decision.3 says the Autonomy watchdog fires "once per partition" and that
recovery is an explicit ISA-88 Restart rather than a silent resume when the
heartbeat returns. HeartbeatWatchdog.fired
(internal/adapter/heartbeat.go) implements that, and it is correct. What was
wrong is who cleared it.
Watchdog().Release() is the only thing anywhere that clears the latch. It is
reached from POST /api/v1/hold/release, whose only caller in the product is
PhaseReconciler.releaseEdgeHold, and that call used to require
Phase.status.edgeSelfHeld. That field is set in exactly one place —
reconcileEdgeSelfHold, whose evidence is an HTTP GET to the runtime pod. So
the release was gated on the control plane having watched the edge hold
happen, and on a partition it has not: the pod is on the far side of the fault.
The two clocks make that the ordinary case rather than a race. The control
plane's own road to Held is reconcileUnitFaultHold, which fires at
pod-not-ready plus a 30-second grace and sets no edge flag. The watchdog fires
at holdGraceSeconds, 60 seconds by default, behind the partition. The loud
road reaches Held first, every time. The Restart that follows found the gate
false, sent no release, and left the latch set for the life of the runtime
process.
Measured on the bench on 2026-08-25, two reps of the same partition drill on one
runtime process. Rep 1 did everything this ADR promises: the watchdog fired at
T0 plus 58.1 seconds, ran the unit's safeStateChart, and settled 0.3 seconds
later. Rep 2 ran nothing — no watchdog line, no hold, no hold event on the
edge's disk, no AuditRecord — and both outputs read HELD_LAST_VALUE for the
whole partition while the unit still reported
runtimeBinding.edgeHold.baselineArmed: true with program: unit-baseline and
dcs_runtime_hold_armed still read 1.
Three things follow.
The command out of Held is the whole condition. releaseEdgeHold is sent
on any ISA-88 command that leaves Held — Restart, Stop or Abort — and the
edgeSelfHeld status fields are cleared only when this phase was carrying them.
The alternative considered was to make the flag honest instead, by having
reconcileUnitFaultHold poll the edge before reflecting. That is more faithful
to what the field means and it was rejected as the fix, because it leaves the
release depending on an observation that can still fail, and the failure is
silent. The observation is worth having; it is not worth gating safety on.
A no-op release costs a request, and a missed one costs the next partition.
handleHoldRelease clears the self-held state and the latch unconditionally,
and a runtime that is not self-held has nothing to clear, so the request is
cheap and idempotent. A command out of Held is not a hot path. This is the
same trade §Decision.3 already makes in the other direction: fire once and make
the operator ask for the resume.
Armed and able-to-fire are different facts, so they get different signals.
Nothing in the product said the watchdog had spent itself. hold_armed,
holdBaselineArmed and edgeHold.program are all about the PROGRAM, and all
three answer the same in both states, because the chart really is armed and
simply cannot be run. dcs_runtime_hold_watchdog_latched and
Unit.status.runtimeBinding.edgeHold.watchdogLatched are the discriminator,
reported on the heartbeat alongside the arming. The latch gauge is published
only where a watchdog exists, so an absent series means "this unit has no edge
watchdog" and 0 means "it has one and it can fire" — the #1648 rule, that a
signal's absence must mean exactly one thing.
The status field obeys the same rule by carrying no omitempty. The gauge
gets its absence-means-something from the mode, and a bool on a CRD has no such
luxury: omitempty erases a false, and an erased false is byte-identical to
what a control plane predating the field writes. A reader would have to settle
which one it was looking at from the build, which is the confusion this field
was added to end. So the value is always serialised, exactly as baselineArmed
beside it is, and the schema stays +optional because an object stored before
this amendment really does lack the key.
TestEdgeHoldArmingWritesTheLatchWhenItIsFalse holds it, because re-adding
omitempty compiles and passes everything else.
TestReconcile_Restart_ReleasesEdgeHold_AfterUnitFaultRoad drives Held by the
unit-fault road and fails against the old gate.
TestReconcile_NoReleaseWithoutACommandOutOfHeld is what stops it passing for
the wrong reason: a release sent on every reconcile would satisfy the first test
and would also clear a latch behind a partition that is still on, which is the
silent resume this ADR forbids.
Amendment (2026-08-25, #1833): the arming reaches the surfaces a person opens¶
The amendment above records the last report on
Unit.status.runtimeBinding.edgeHold, and the amendment before it adds the
field that separates armed from able-to-fire. Neither put the answer anywhere a
person could read it. grep -rl EdgeHold over the tree returned the adapter,
the two controllers, the API types and their tests, and no file under
internal/gateway/ or cmd/dcs/. So the question this whole ADR exists to
answer — is this equipment covered right now — was reachable through kubectl
get unit -o jsonpath and nothing else, four months after the record was
written.
Placement was a founder call, taken on 2026-08-25 against measured mock-ups drawn in the shipped gateway CSS. Three surfaces carry it, and the split between them is which failure each one survives.
The Unit detail's Properties block carries an Edge Hold row, reading the control plane's record. It is the one that survives the partition the arming exists for, and it costs 42px of a 316px block. The row renders on every unit, healthy ones included, which is the opposite of the rule the three #1772 neighbours follow: those carry exceptions, so a permanent line saying nothing is wrong would be annunciation with no decision behind it (ADR 0029). This row exists so a commissioning engineer can confirm the promise holds, which a row that appears only when it is broken cannot do. Runtime Ready three rows above it is green on every healthy unit for the same reason.
dcs get runtime carries three lines, read live off the runtime pod's own
HoldController and Watchdog through the diagnostics proxy. It is the
freshest reading and the one that stops existing when the pod does. Where it
disagrees with the row above, the disagreement is the finding.
dcs get units carries an EDGE HOLD column, off the same control-plane
record. Neither of the other two closes the gap the issue opened with: one is a
gateway page and the other reaches the pod, which on a partition is on the far
side of the fault. A drill script and an engineer at a terminal have this.
The verdict is computed once, and one of its inputs is not on the report.
Six readings live in edgeHold plus the unit's own spec.safeStateChart, and
two of them are byte-identical on the wire. A report saying nothing is armed
means a chart failed to arm on a unit that declares one, and means the
configured posture on a unit that declares none — opposite remedies, and the
second is not a fault at all. unitEdgeHoldToDTO is the one place that
decides; the browser and the CLI render its summary and never re-derive it.
dcs get runtime is the exception and has to be: the runtime holds no Unit
spec, so its Armed Program line says what is armed and deliberately does not
say whether that is wrong.
The latch outranks everything the report says about the program. On the #1818
bench rep the program was named and the baseline slot was armed while nothing
either of them described could run. Ranking watchdogLatched below
armed renders that unit green, which is the reading that lost a whole second
partition.
TestLatchedOutranksAnArmedProgram fails against a latch demoted below the
program fields, and TestNothingArmedIsDecidedByTheDECLARATION fails against a
verdict that stops reading spec.safeStateChart.
Amendment (2026-09-02, #1919): a hold is answered by the work that moves on, not only by a Restart¶
The #1804 amendment drew a line: a hold this runtime ran is released by an ISA-88 Restart and by nothing else, and only a hold it found on disk is retired on evidence of presence. The line is right about which evidence clears a live hold — a lease grant or a heartbeat proves the control plane is back, not that anyone has recovered the unit, and a hold sitting under a Held phase looks exactly the same the moment before its Restart as it would if nobody ever sent one. It was wrong about the premise underneath it, which is that every hold has a Restart coming.
Measured on the bench on 2026-09-02, drill 12 rep c of the #1894 verification,
chart 0.7.1. The unit fenced for 9.5 s (#1909). The runtime ran its bounded
hold, dcs_runtime_self_held went to 1 at 20:56:28.7, the lease was re-granted
and the fence cleared at 20:56:38, and the gauge stayed at 1. The batch that was
fenced had already completed before the cord came out, so the hold ran on an
idle unit under the baseline. No phase was Held by it, so no Restart was ever
going to arrive for it, and the next batch's Start is not a Restart. Two full
batches then ran to Complete on that runtime, both regulating at setpoint, and
the gauge still read 1 at 22:40. Rep 2's harness row carried
self_held=min=1 max=1 through a rep in which nothing held, which is the cost:
the signal could not have seen a second hold if one came.
What answers a hold. Two more roads, and both are the control plane's word
on the same POST /api/v1/hold/release, which now carries the road in its body
and records it on the released audit event:
later-phase-started— a phase starts on the unit and finds the runtime self-held under work that is not this phase's. A unit that takes a fresh Start has been recovered by that Start. The discriminator isedgeSelfHoldIsThisPhases, the reading the reflection already uses: a phase-scoped program names its phase, and a baseline hold that triggered before this phase started is not this phase's. A hold that is this phase's — the runtime fenced mid-Running and the ActionChart is re-entering after the outage — is left standing for the Restart or for the phase to end.engaged-phase-ended— the engaged phase reaches Complete, Aborted, Stopped or Idle. Terminal is the other end of engaged, and there is nothing left to Restart. The release is unconditional on the transition, on the argument #1818 already made for the Restart's: one request on a transition that is not a hot path, and a no-op on a runtime that is not self-held.
The Start road reads GET /api/v1/status rather than riding the arm request,
because the bench's phases declare no holdingChart and the arm request never
leaves for a phase without one. A road that could not fire on the unit the
defect was measured on would not have closed it. The read is one GET on a phase
start; when it fails the hold is left for the next road rather than guessed at.
The watchdog latch shares the premise, and it is the safety half. #1818
made the Restart's release clear HeartbeatWatchdog.fired, on the same
reasoning: recovery from an edge hold is a Restart. In Autonomy the shape
measured here — a baseline hold on an idle unit, then a fresh batch — leaves
the watchdog latched under a live batch, so the next partition runs no chart
at all while every instrument reports the program armed. That is worse than
the gauge. Both hold-side states are released together on every road; Failover
has no watchdog and is unaffected on that half.
What is deliberately not done. The runtime does not clear itself on
presence: RetireInheritedHold is unchanged and still refuses a hold this
process ran, because presence cannot tell a hold awaiting its Restart from one
nobody will Restart. The runtime refuses a release naming a road it does not
know (400), so a misspelling in the control plane cannot write a road that does
not exist into the audit trail. An absent body or an empty reason is the
Restart, which is what every caller before this amendment meant. During a
mixed-version rollout an older runtime ignores the body and records every road
as control-plane-restart; an older operator's bodiless release lands as the
Restart on a newer runtime. Neither leaves a hold standing.
The cost is the reflection's own: the Start road compares the runtime's
triggeredAt with the phase's startTime, two clocks, and it abstains rather
than refuses where the reading is unavailable. A hold triggered within clock
skew of a phase's Start reads as that phase's own and waits for the terminal
road, which is the fail-safe direction.
TestReconcile_Start_AnswersAHoldThatIsNotThisPhases drives the bench shape
and fails against a tree with either road removed;
TestReconcile_Start_LeavesThisPhasesOwnHoldForTheTerminalRoad is what keeps
the Start road from answering a hold the phase is still under.
TestHoldRelease_NamesTheRoadTheControlPlaneTook holds the runtime to
recording the road and releasing the latch with it, and
TestHoldRelease_RefusesARoadThatDoesNotExist holds the refusal.
Amendment (2026-09-03, #1928): the audit bridge is subscribed by the leader, and every replica connects under its own id¶
Context. The chart has run the physical-operator at two replicas
since #1893. Both replicas built their MQTT client at process start under
the literal id physical-operator, and both subscribed the hold-lifecycle
and safe-stop bridges before leadership was known. MQTT requires a client id to be
unique per session, and Mosquitto enforces it by closing the older session
when a CONNECT arrives under its id. The bench broker logged 495 New client
connected lines in 70 minutes, alternating between the two pod IPs, with
already connected, closing old connection before each. The leader's
subscription to the hold-event topic was dropped and remade at that cadence,
and a hold event published into one of those gaps reached nobody.
Decision.
- The client id is the pod's. Every operator names its session after
POD_NAME(the downward API, set by the chart) and falls back to the hostname, which the kubelet sets to the pod name. The bare component name never reaches the broker: with neither, the id carries a random suffix. - Every replica connects at start. The connection is not the leader's, because a replica taking over needs a session already up: its first unit-state announcement must not wait on a TCP, TLS and CONNECT round trip inside the grantor gap (#1909). The standby's connection is idle.
- The subscriptions are the leader's. The two bridges are subscribed by a
leader-election Runnable (
AuditBridgeSubscriber) on winning the election and unsubscribed when the manager stops. A subscription is a writer: each event becomes anAuditRecordunder a generated name, and two replicas holding it would write every hold and every safe stop twice. Before this amendment the trail stayed single only by the accident this amendment removes, because whichever replica held the connection at that instant was the one materialising. A standing-down leader (#1894) keeps running after its manager stops, and the unsubscribe on stop is what keeps it from writing beside its replacement.
What this does not change. An event published while no leader holds the subscription is received by nobody and kept by nothing, and the election gap is one instance of that. The runtime's queue (#1779) survives a lost broker on the publisher's side; it does nothing for a subscriber that was not there. That is #1927, and a stable session the broker could resume for the leader role is not its answer here, because a stable id is exactly what two live processes cannot share.
Consequences. The broker log shows one New client connected per pod
start and no already connected lines. dcs_mqtt_connected{component=
"physical-operator"} is read per pod through Prometheus's own pod label;
the standby's reading is its idle connection, which is true. The chart gate
liveness-grantor.sh asserts POD_NAME comes from the downward API and not
from a value in the template, and TestMQTTClientConfig_TwoReplicasCarryTwoClientIDs
holds the binary to it.
Amendment (2026-09-03, #1927): the durable session the leader alone holds is the reader that outlasts the absence¶
The #1779 amendment made the runtime's half of the edge audit path a disk write, so a hold never waits on the network to reach its safe state. The runtime's half held. The record still did not survive.
Measured on the bench on 2026-09-03, chart 0.7.3, the first grantor-absent
rep. The physical-operator was scaled to 0 for 150 s with an idle
Failover unit. The runtime did what ADR 0006 asks — bounded edge-local
hold, settled into safe state, self-fenced, fenced for 132 s — and published
every one of those events to a broker that was up throughout. The trail
records none of them. It records only the released, 37 s after the operator
was back, so it says a hold was released that it never says ran.
The operator is the only thing that materializes these events into
AuditRecords (the runtime never writes the apiserver). It had subscribed
with a session that ended when its connection did, so the broker held nothing
for it between connections, and the events published while it was absent went
to a broker with no subscriber and were gone. This is not the #1780
store-and-forward gap: the broker was up, the runtime published, its queue had
nothing to hold. The loss is on the subscribing side, and it falls on exactly
the fence that matters most — the one shape that fences a Failover unit is
the grantor going away (#1893), and every fence of that class runs while the
only reader of its record is the absent operator.
The fix is a durable subscriber. The two audit bridges — the hold
lifecycle and the terminal-stop safe state (#1283) — move onto their own MQTT
connection that presents a stable role client id (physical-operator-audit-bridge)
with a week-long SessionExpiryInterval and clean-start false. The broker
keeps the session, and every QoS 1 message published to it while the operator
is away, until the next leader reconnects under the same id and drains it.
Three things make that connection what it is, and each is load-bearing:
- It is the leader's alone (
NeedLeaderElection). A session is keyed on the client id, so two live connections under it are a takeover, not two subscribers, and the standby reconciles nothing to give a record to. - Its id is a role, not a process. The pod that comes back after an eviction has a new name, and a session keyed on the pod name would be orphaned by the very restart it exists to survive. This is the opposite of what the publisher on the same operator needs, where an id per process is what stops two replicas taking each other over (#1928) — which is why the audit bridge is a second connection and not a setting on the first.
- Its handlers are registered before the connection starts. A resumed
session's queued messages arrive on the heels of the CONNACK, and paho drops
an inbound publish that no registered handler matches, so a handler
registered after
Connectreturned could miss the replay the session was kept for.
The broker's own max_queued_messages bounds how many messages it will hold
for that session; the chart raises it from Mosquitto's default of 1000, since
a grantor loss fences every Failover unit at once and a hold is a dozen
records each.
The cost is at-least-once delivery on a resumed session: a message this pod
acknowledged to the apiserver but not to the broker when it died is
redelivered to its successor and becomes a second AuditRecord. The recorder
already takes that trade on its own retries, where a duplicate is "strictly
preferable to a drop," and it is the same trade here.
Between the runtime's store-and-forward queue (the broker being away) and this durable session (the subscriber being away), the record now has a home at every point on the path.
TestIntegration_ASessionOutlivesItsSubscriber drives the bench shape against
the pinned broker. A subscriber leaves, a publish lands while it is gone, and a
new process under the same id comes back. It runs both arms, so the durable
session's delivery is proven against a plain session's loss.
TestAuditBridgeSubscriber_SubscribesBeforeItConnectsAndDisconnectsOnStop
holds the register-before-connect order and the disconnect-not-unsubscribe on
losing leadership, and NeedLeaderElection keeps it the leader's alone.
TestAuditBridgeClientConfig_IsADurableSubscriber holds the session and the
role id. TestAutopahoConfig_CarriesTheSession holds that the session fields
reach the wire and that a client asking for no session is handed none. The
annunciation half, that the same absence raised no alarm for the fence,
belongs to #1890 and is recorded there.
Amendment (2026-09-11, #2137): a report belongs to the runtime that made it, and the record follows the report¶
The #1806 amendment put the runtime's last arming report on
Unit.status.runtimeBinding.edgeHold, and the #1833 amendment put that record
in front of a person. Two things about the record were left to the reconcile
loop, and a manual failover of an Autonomy unit found both.
The first is whose report it is. The heartbeat loop keeps one memory per
unit: what the runtime last answered, and which chart this loop landed on it.
A re-bind moves the loop onto a different pod at a different address and the
memory came with it. performFailover rebuilds the binding with no record,
and the next pass copied the loop's memory of the previous runtime into it,
stamped now, as the promoted runtime's own answer. On the multi-node capture
stack that read armed for the 21 s between the re-bind and the broker's arm
landing, over a runtime that had nothing armed. On the bench the previous
runtime's last decoded answer was the one taken before its own arm landed, so
the same copy read nothing armed, and the loop's memory that it had already
armed that chart kept the optimistic arm from being posted to the new pod. A
runtime's arming is process memory. Nothing learned from one process holds for
the next, so a re-target onto a different address drops the report, the arming
memory and the last-acknowledged instant together, in both brokers. The
in-place road, a pod that came back at another address under the same binding,
clears the record on the Unit for the same reason it already clears
leaseEstablishedBy.
The second is when the record moves. It is written only when the reconciler
runs, and an Idle unit has no requeue of its own, so a beat that learned the
promoted runtime was armed waited on whatever next happened to reconcile the
unit. Both brokers now nudge the reconciler when the answer moves, on the same
channel a lease transition already uses. And the answer moves sooner: a
re-target and the pod's readiness edge both arm and beat at once rather than at
the next tick, which is a third of the grace away, and a beat that lands an arm
asks once more so the record follows the act by seconds. The window an
operator saw nothing armed over a standby that was armed is the tick plus the
next reconcile; it is now the round trip.
TestAManualFailoverRearmsTheBaselineOnThePromotedRuntime drives the shipped
road through the re-bind on both transports the channel has, HTTP/1.1 and
HTTP/2 over TLS with the operator's own transport shape, and holds the three
claims apart: the record says nothing until the promoted runtime answers, the
arm lands inside one beat interval of the readiness edge, and the reconciler is
nudged when the report moves. Each was proved by breaking what it guards.
Amendment (2026-09-22, #2155): a marker the phase has already answered is the hold being recovered from¶
The #1723 road to Held reads Unit.status.conditions[HeldByRuntimeFault] and
reflects it onto the phase running on the unit, and the #1818 amendment settled
that the command out of Held is what answers a hold. The two met on the bench on
2026-09-22 at 17:51:40Z, on the cord-pull rep of the manual-failover take. A
confirmed-fenced failover of bench-loop-01 landed at epoch 122, the runtime
came Ready on the standby, and the drill issued the documented Restart to the
batch ten seconds later. Within one second the physical operator took the
Restart (Held → Restarting → Running, marker removed), the procedural operator
took the phase's (Held → Restarting), and then reflected the unit-fault hold
onto the phase twice, and the tree held with it (#1823). The batch sat Held and
the drill aborted it after 90 s. The same recovery had succeeded on every
reboot-road pass that hour.
The physical layer removes the marker when the Unit leaves Held, in the same status write as the state. The phase controller is a different process reading the Unit through its own informer, and its post-Restart pass ran before that write reached its cache. So it read the marker the Restart had just answered as a fresh fault. The race is by construction, and whether it is lost is watch latency.
Three things follow.
The reflection records which marker it answered.
Phase.status.reflectedUnitHoldAt is the marker's own LastTransitionTime,
written by reconcileUnitFaultHold when it drives Held. runtimeFaultHold,
the one lookup both the reconciler and the SFC status publisher read, refuses a
marker stamped at or before it. A new fault always arrives with a later stamp,
because the physical layer removes the condition when the Unit leaves Held and
a condition that was absent is stamped afresh when it is set again. So the
record cannot mask one.
The anchor is the marker reflected, not the moment the phase left Held. The
issue proposed the latter, and it was rejected because of a road it breaks. A
phase Held for another reason (a sync barrier, a coordination wait, an operator
Hold) while the fault arrived never reflected the marker. By that anchor its
resume would refuse the marker as older than the Restart and run the phase over
held equipment. TestPhaseReconcile_RestartOutOfAnotherHold_StillReflectsTheFault
is that road. The record is cleared by Reset, because the next Start is a fresh
run against whatever the equipment says then.
The UnitProcedure's barrier reads a stamp of its own, the other way round.
convergeUnitState leaves a Unit the physical layer holds where it is, awaiting
the explicit Restart. That Restart is the UnitProcedure's own command out of
Held, now recorded in status.leftHeldAt. A Unit still Held under a marker
stamped strictly before it is a Unit the Restart never reached, which is the
lost command the #1151 repair exists for, and the repair sends it again. The
comparison is strict because both stamps carry second precision, and the safe
reading of a same-second marker is a new fault. This read was already live
rather than cached (#1774), so the stale-cache half of the defect never applied
to it. What applied was a barrier that read every marker as awaiting a Restart
that had already been given.