Skip to content

High Availability and Failure Modes

HA Architecture Overview

Cloud-Native DCS builds on Kubernetes-native HA primitives and adds process-control-specific mechanisms:

Mechanism Component How It Works
Leader election All operators controller-runtime Lease; only one replica reconciles at a time
State recovery All controllers ISA-88 state machine rebuilt from CRD .status.state on every reconcile
FB program persistence Unit runtime Every deployed program saved to hostPath (networks/all.json) on deploy and stop; all of them redeployed from disk on startup, before the control plane is reached (ADR 0077, #1748)
Driver reconnection Unit runtime Exponential backoff (100ms-30s) with proactive health monitor every 5s
Grace-period health check Unit runtime The gRPC READINESS service reports NOT_SERVING after all FIELD drivers have been disconnected for 60s. The liveness service is a different question and stays SERVING through it, because a restart reconnects no cable (#1918). The count excludes simulation drivers, which report connected for the life of the runtime and so cannot say anything about a plant; a unit with no field drivers stays SERVING throughout (ADR 0075, #1743)
Watchdog Hold Unit controller Queries runtime status API; issues ISA-88 Hold when ANY FIELD driver is down during an active batch — Critical when every one of them has gone, High when some are still answering (ADR 0078, #1754). A simulation driver is never counted and a unit with none abstains (ADR 0075, #1743). Keeps reading driver health across the Hold it issued, and clears its own alarm and degraded condition when the drivers reconnect, whatever the operator does with the batch (#1735)
Degraded-program alarm Physical operator Raises a High equipment alarm naming every block that faulted on the runtime's last scan, from ControlProgram.status.faultedBlocks, and clears it when the scan comes clean (ADR 0078, #1754). No batch effect: a degraded program is still regulating every healthy output (ADR 0009). It is what annunciates a field WRITE that failed, which reached health, metrics and the HMI and no alarm at all
Control lease + self-fence Unit runtime (availability mode Failover) Runtime holds an operator-brokered lease; on expiry it self-fences (stops FB output writes, reads continue) on its own clock — works under control-plane partition (ADR 0006)
Hold-then-resume failover Unit controller (availability mode Failover) Re-binds a unit's runtime to a designated standby node after lease expiry + safety margin; manual fenced failover for Autonomy (ADR 0006). The high-severity Alarm it raises says the runtime is not back to normal, and clears when the replacement is Ready and holding its control lease on the standby, whatever the operator does with the batch (#1733)
Edge-local holding Unit runtime On a control-plane partition (heartbeat watchdog fires) an embedded SFC engine runs the armed hold chart — the active phase's holdingChart, else the UnitSpec.safeStateChart baseline — to drive a deliberate sequenced safe state by writing the unit's own control-module tags; the FB scan stays the sole I/O writer (ADR 0008). A holdingChart that addresses another unit's tags, or calls a builtin needing an operator, the apiserver or another machine, is refused when the runtime is asked to stage it, so the baseline covers that phase and the operator learns which at phase start (#1406). The baseline is armed for every unit in either availability mode, and re-armed whenever the runtime reports its baseline slot empty, so a restarted runtime pod recovers the posture on the next lease renewal or heartbeat; Unit.status.runtimeBinding.edgeHold carries the last report (#1806). The Autonomy watchdog fires once per partition and then latches, because ADR 0008 requires recovery to be an explicit ISA-88 Restart. A returning heartbeat must not resume the unit silently. Any command out of Held releases the latch. Until #1818 the release was sent only when the control plane had watched the edge hold happen, which on a partition it usually has not, so the second partition on one runtime process ran nothing. edgeHold.watchdogLatched and dcs_runtime_hold_watchdog_latched are what say the watchdog has spent itself; the arming fields cannot, because the chart really is armed and simply cannot be run
Link redundancy (NIC bond) Controller node (deployment-configured) Active-backup bond per network zone; a cable, NIC, or switch-port fault fails over below the runtime and never surfaces as a partition. Members on separate switches where available — though single-homed remote I/O caps what that buys for the field zone (ADR 0026). Deployment-layer machine config, transparent to the FB scan and drivers
MQTT store-and-forward Unit runtime Buffers up to 10K messages during broker disconnection
Liveness/readiness probes All components Kubernetes restarts unhealthy pods; removes unready pods from service
Startup probes All operators, gateway, historian Allows up to 150s for initialization before liveness probe takes over
PodDisruptionBudget All deployments Auto-enabled when replicas > 1; prevents voluntary disruption from removing all replicas
Pod anti-affinity All deployments Spreads replicas across nodes when replicas > 1

Failure Mode Analysis

The coordination layer is elastic where that is safe, and deterministic where it must be. Batch state lives in the model (the Batch and Procedure resources) and never in an operator pod, so the pods above the process are ordinary replaceable Kubernetes workloads. One of them is not, and the exception belongs in front of the table. The physical-operator is where both edge liveness brokers run, so its absence is the one that reaches the field. A Failover unit self-fences once its control lease stops being renewed, and an Autonomy unit runs its armed hold chart once the heartbeat stops arriving. That is the design working as specified. It is also why the chart gives this one operator a standby by default (#1893). The unit runtime scanning real I/O is pinned to its device node and never disturbed. The clip below demonstrates the first row of the table live. The procedural-operator pod (the SFC engine) is deleted mid-batch, Kubernetes replaces it, and the replacement resumes the chart exactly where the model says. The unit-runtime pods show zero restarts throughout.

Killing the SFC engine mid-batch: the procedural-operator pod is deleted while a batch runs, Kubernetes reschedules it, and the chart advances charge → inoculate under the replacement. Unit runtimes untouched, RESTARTS 0.
Failure Mode Detection Response Recovery Time Impact
Operator pod crash Kubernetes liveness probe (20s period) Pod restart (RestartAlways) + leader election handoff 15-30s for a pod delete, which is the fault this row describes. Not bench-validated. The figure is the leader-election lease expiring and an immediately-rescheduled replica taking it. It is derived from the lease durations, and nothing has measured it. A node loss is a different clock and this figure does not cover it: a pod on a NotReady node is not evicted until its node.kubernetes.io/unreachable:NoExecute toleration expires, 300s by Kubernetes default, and a single-replica Deployment has nothing to hand the lease to until that eviction produces a pod. A handoff measured in the low hundreds of seconds after a node loss therefore names the eviction timeout. Leader election is not the term being measured there. The bench measurement is #1888, which records the replica count beside every figure so the two cannot be confused. Since #1893 the chart defaults the physical-operator to two replicas, so on that one the eviction clock no longer governs and the figure is the leader handoff again. Measured 2026-09-02 on chart 0.7.1 (drill 12 rep c, the leader's own node cut): the standby acquired at T0+16s and the replacement grantor was created at T0+101s and Ready at T0+111s, against T0+397s for the single-replica procedural-operator on the same rep. The unit still fenced, for 9.5s. The standby was renewing at T0+21s, and the runtime's deadline after a renewal is the lease minus the renewal interval, 20s and not 30s, because the operator anchors every renewal on the age of its previous acknowledged one so that its own Expired decision and the runtime's self-fence land on the same instant. The grantor gap had to fit inside 10 to 20s and did not. That is #1909, and the fix is shipped and not yet bench-verified (ADR 0006, 2026-09-02 amendment): the physical-operator elects at 6s/4s/1s (physicalOperator.leaderElection), where controller-runtime's scaffold default is 15s/10s/2s. It starts every liveness loop from its cache the instant it wins, before the reconciler has reached any Unit. And it renews at a sixth of the lease, where a third was the cadence, so the budget after a renewal is 20 to 25s and the worst handoff about 10s. The anchor did not move, because it is what keeps the operator's expiry and the runtime's fence on one instant. The rep that proves it was taken on 2026-09-03 on chart 0.7.4, the same rep c aimed the same way with the standby pinned to a node the cut did not touch: the standby acquired the lease at T0+2s, started its renewal loop 1s later and warm-started edge liveness from its cache in 865ms, and dcs_runtime_fenced read 0 across 106 scrapes with no lease expiry recorded at all. Measured 2026-09-03 on chart 0.7.3 (the same rep c, with the standby pinned to the cut node's apiserver): the election and the warm start were both on time, acquisition within 6s of the standby having an apiserver and 1.4s to the first renewal loop, and the unit fenced for 29s anyway. The standby had no apiserver for 43s. Its HTTP/2 connection was pinned to the cut node, a node losing power closes nothing, and client-go keeps using such a connection until its health check gives up, 30s read-idle plus a 15s ping timeout by default. That is #1930, fixed and bench-verified (ADR 0006, 2026-09-03 amendment): the physical-operator sets client-go's two timeouts to 2s each (physicalOperator.apiserverHealthCheck), so a dead apiserver connection is given up inside 4s of its last frame. The detection runs in front of the election and does not overlap it, so the worst handoff is about 14s against the 20s budget. A fresh dial can still land on the dead endpoint for the ~25s until the apiservers prune it, one time in three per dial on three control-plane nodes, and each such dial costs 2s of the 5s margin. Measured 2026-09-03 on chart 0.7.5 (drill 12, the same rep c with the standby on a third node pinned to the cut node's own apiserver, and the VIP deliberately elsewhere so its failover could not mask the pin): the standby logged http2: client connection lost at T0+1.8s where the 0.7.3 rep waited 43s for that line, acquired the lease at T0+10.0s, and warm-started edge liveness in 674.6ms. Nothing fenced, no lease expired, nothing re-bound, and the AO held last value on both channels. The 10s sits inside the 20s a runtime has after a renewal in the worst phase, and T0 skew of 9s makes it a lower bound. The two single-replica operators on the cut node moved at T0+368s in the same rep, which is the eviction clock this row distinguishes and not leader election. Measured 2026-09-01, four reps, on chart 0.6.3 — which predates both that change and #1890's (cndcs-deploy-bench § Drill 12). All three clocks this row distinguishes were seen, and they are far apart. A graceful drain of the node carrying all four leases moved all four within ±3s, which is leader election alone and is what the 15-30s prediction was reaching for. A node loss with the replica on it left the lease held for 372s at replicas=1 (2026-08-31). That is the NoExecute eviction timeout and not leader election, exactly as this row says. And a third case the row did not anticipate: an abrupt loss of the node holding the API VIP moved no lease at all, because the operators were on a different, healthy node and all three exited 1 with leader election lost when the API went away for 57s. They were replaced in place, and the drill gives a same-node replacement its own verdict, so it is never counted as a handoff. #1893's replica count covers that third case only by chance: the Service address is shared, but each pod's connection is pinned to one apiserver endpoint, so the follower rides the outage through when it is pinned elsewhere and dies with the leader when it is not (measured 2026-09-02: two operators on one surviving node, opposite fates). When both are pinned to the lost endpoint the plant has no grantor for the length of the API outage. That is #1894, and on the bench it expired a unit's 30s lease and fenced it for 37s. Fixed: the grantor no longer leaves with the manager. A physical-operator whose manager stops without being asked to stands down for physicalOperator.leaseStandDown (default 90s) before it exits, going on renewing leases and heartbeats it already holds while reconciling nothing, because a grant of liveness travels to a runtime pod IP and never touches the apiserver (ADR 0006, 2026-09-01 amendment). The price is paid by whichever process takes over: an operator that has just acquired leadership refuses to read a lease expiry as a fence until the predecessor's window, the lease duration and the re-bind margin have all elapsed, so a partial partition cannot re-bind onto a runtime the old leader is still holding alive. A human-certified fence does not wait. The graceful drain in the same drill fenced for 6.4s with the apiserver never away, and that half is answered separately: the operator now sets LeaderElectionReleaseOnCancel, so a leader stopped on purpose hands its lease straight back. It no longer holds it for the whole LeaseDuration, and the standby's broker starts renewing unit leases with the 30s budget still mostly intact. The abrupt half is bench-verified, 2026-09-02, chart 0.7.1 (drill 12 rep 2): the leader lost its pinned apiserver, stood down at T0+14s, no lease expired, no unit fenced, and the incoming leader was renewing 6s before the outgoing one let go. The window clears a ceiling and not a sample: a Talos L2 VIP failover is bounded at 60s by the etcd session TTL that elects the holder (#1905), and 90s clears it by 30s. A loss of quorum is bounded by nothing and is not what the window is for: one on the same day kept the API away 307s and fenced the unit for 165s at the end of the window, with the grantor still reaching it. The graceful half and the plant cost were measured the same evening, on a 420s cut phase built so the cord could come out with the loop live (cndcs-deploy-bench 29e0967). A graceful drain of the node carrying the leader and the VIP moved the grantor lease at T0+1s, expired nothing and fenced nothing, with the AO held and the PV band at 59.998 throughout. An abrupt cut of the node carrying the VIP and the leader's pinned apiserver, leader on a third node, engaged the stand-down at T0+15s, had the standby renewing at T0+24s, the API back at T0+59.0s, and again expired nothing and fenced nothing, AO held at 19628 and PV in band for the whole window. The cost of the fault to the plant was zero on both. The same rep found that the single-replica procedural-operator, coming back from eviction at T0+424s, aborted the running batch (#1917), which is a defect in that operator and not in the grantor. Fixed, and bench-verified on 2026-09-03 on chart 0.7.4, where the returning operator's converger logged its two Converging Unit state lines for the batch's own UnitProcedure and for none of the 336 stale ones, sent no Abort, and let the batch complete: the restarted process reconciled every terminal UnitProcedure in the namespace off its informer's initial list, and the unit-state converger sent each one's command to the unit without asking whether that UnitProcedure was the one running on it, so the last stale Aborted one to land aborted the live batch. The converger now acts only for the element the Unit's own status.activeWork names, or, when that is empty, for the batch the unit is allocated to. Rep 3's log shows a fenced runtime failing its liveness probe and being restarted about 60s into the fence (#1918) Reconciliation paused; running batches continue autonomously via runtime — for the procedural, batch, control and alarm operators. The physical-operator is not one of those. It runs both edge liveness brokers, the ADR 0006 control lease and the ADR 0008 Autonomy heartbeat, and it runs them in the leader process and in no other. A runtime cannot tell a grantor that has died from one it is partitioned from, and it must not, so while that process is absent a Failover unit runs its bounded hold and self-fences at availability.leaseDurationSeconds (30s unset) and an Autonomy unit runs its armed hold chart at holdGraceSeconds (60s unset) and then latches until an ISA-88 Restart. One control-plane node losing power left the plant with no grantor for 372s on the bench, against those two budgets (#1890, 2026-08-31, n=1). The chart answers that with the default replica count above, preferred pod anti-affinity, a PodDisruptionBudget that leaves one replica standing, and a NoExecute toleration well under Kubernetes' 300s; deploy/helm/cloud-native-dcs/tests/liveness-grantor.sh holds all four, each against the render that removes it. On the bench that turned 372s into a 16s handoff and a 101s return to redundancy, and left a fence of about ten seconds that was the renewal arithmetic and not the chart (#1909, fixed and bench-verified on 2026-09-03). The apiserver going away is a different fault with a different clock, and the chart answers it separately with physicalOperator.leaseStandDown, held by deploy/helm/cloud-native-dcs/tests/lease-stand-down.sh (#1894). Dropping physicalOperator.replicas back to 1 is supported and restores the exposure in full
Unit runtime pod crash Kubernetes liveness probe Pod restart + FB network replay from disk 20-40s I/O ceases during restart; batch Held by watchdog if enabled. The Hold names what failed (#2153): its HeldByRuntimeFault reason is PodNotReady when the pod stopped answering on a node that is still Ready, and NodeNotReady when the node itself stopped reporting and the pod's condition merely followed it, which is what a partition or a lost machine produces. The Unit page's hold-provenance chip reads the matching sentence, so an operator is sent to the pod or to the machine and not to the wrong one
All I/O drivers disconnected gRPC readiness (60s grace) + watchdog Watchdog issues Hold; pod goes NotReady after the grace period and is NOT restarted. Until #1918 one health verdict answered both probes and the kubelet restarted the pod about a minute after the grace, which reconnected no cable and threw away the held state. Liveness now answers only whether the process is alive and scanning 60s grace to NotReady. The Hold is bench-validated (n=1, #942 drill 2a); the restart it used to record never was. Cutting a unit's whole field path Held it at +21.1s, well ahead of the grace period, on one Critical alarm naming every driver. The alarm cleared itself to ClearedUnacknowledged when the path came back, which is the #1735 clearing edge on hardware Batch placed on Hold; manual Restart after drivers recover
Single driver disconnected ReconnectingDriver health monitor (5s) Auto-reconnect with exponential backoff 100ms-30s Partial I/O loss; affected channels return stale values
MQTT broker down Store-and-forward queue + reconnect; dcs_mqtt_connected == 0 on operators/gateway Runtime buffers messages and replays on reconnection; operators and gateway retry in the background (capped exponential backoff, ≤30s) and resume publishing automatically — including when the broker first comes up after them. A runtime that starts during the outage comes up the same way since #1933: its connection is opened in the background, its health and lease endpoints open regardless, and a Failover runtime is granted its lease with the broker still away. Before that fix a starting runtime waited inside Start for the broker with no timeout and listened on nothing, which is what sustained the 23-minute re-bind storm on the bench on 2026-08-31 (ADR 0006, 2026-09-03 amendment). Fixed and not yet bench-verified Transparent Historian data gap; HMI shows "No MQTT" until reconnect; no control impact
etcd unavailable Kubernetes API errors Controllers back off; running batches continue Kubernetes-dependent. Not bench-validated, and this row had no drill behind it at all. The #942 campaign takes edge nodes, partitions, bonds and a VLAN quarantine, and every one of those comes out of the control zone. The cold-boot drill (#1349) takes all five machines at once, which removes quorum outright. The question here is whether quorum holds when one member of three goes away. The bench measurement is #1888: one control-plane node powered off with a batch running. The figures are API reachability at the VIP, the front door as a user experiences it, the per-operator leader handoff, the hole in the historian, and whether any Unit is fenced, held or re-bound. The first rep says the last of those was wrong (#1890, 2026-08-31, n=1). One cord out of a control-plane node produced 31 failovers of a Failover-mode unit whose controller and field path were never touched, over 23 minutes, and a 1924.6s hole in the historian. The API VIP was back in 58.9s and the front door in 58.6s, so the control plane recovered in under a minute and the unit went on failing over for another twenty-two. The re-bind is what sustained it: it destroys the runtime whose liveness decides whether to re-bind, and the replacement had one lease duration to start up in. Fixed — the start-up grace and the re-bind budget are the ADR 0006 2026-08-31 amendment — and the drill records the re-bind COUNT now, because a single re-bind reads exactly like a correct response to a dead controller. The storm opened on the replacement operator's own first renewal loop, which is the other thing this row now knows: a leader handoff gives the new process no memory of any runtime having taken its lease, so the binding carries that fact itself (status.runtimeBinding.leaseEstablishedBy) and an established runtime is not read as one that never started. Measured 2026-09-01, four reps (cndcs-deploy-bench drills/ha-drill-runbook.md § Drill 12): cp-3 abrupt, cp-2 abrupt, cp-2 graceful, cp-1 graceful. Quorum holds. etcd members read 3 before and throughout every rep, on a query deliberately asked of a node the rep was not cutting. No member is removed when one goes away and none should be. That instrument is a membership list and not a health check, and on 2026-09-02 it counted a member whose store an earlier abrupt cut had corrupted and whose etcd was crash-looping; the next rep was then the second fault of two and took quorum with it (#1905). The drill now asks each member about its own store and refuses a rep below three of three. The API at the VIP is back in ~58s after an abrupt loss of the node holding it (57.615s, n=238 at 0.5s; 58.934s on the 2026-08-31 rep against a different node; 58.772s on 2026-09-02; 59.001s on the 2026-09-02 evening rep, which is the first one taken with the field live and the first on which that outage cost the plant nothing). The figure has a ceiling: the VIP holder is elected through an etcd concurrency session at etcd's default 60s TTL, so an abrupt loss costs at most what is left of that lease, and a graceful shutdown resigns it, hence 0s (#1905). A quorum loss is not bounded by it: the 2026-09-02 rep that took the second member of three kept the API away 307s and is not interrupted at all when the lost node holds no VIP, or when the VIP holder is drained gracefully (0 failed samples in 7244). The front door at dcs.bench.cndcs.io tracks it within half a second and comes out slightly shorter on every rep. So one member of three going away costs the API up to a minute and costs etcd nothing. It also cost a fenced unit, which is #1894 and is a leader-election consequence and not an etcd one. That one is fixed in the same place the row above records, and bench-verified there. The operator stands down when it loses the API, and the plant keeps its grantor for the length of a VIP failover. The product reads member health since #2112: the cold-boot rep of 2026-09-09 broke cp-1's etcd store, the node came back Ready, and every surface read the cluster healthy for an hour on two members of three. The Servers page, the health verdict, the return readout, the maintenance guard and the server alarms now read which members are serving off the API server's own endpoints, where a control-plane node is listed only while the API server on it can write to its local store, and a member that is Ready with its store dead is graded degraded, alarmed at High with the margin, and refused as a reason to service either of its peers New batches cannot start; running processes unaffected
Network partition (operator-runtime) Watchdog HTTP timeout (5s); lease renewal failure (Failover mode); control-plane heartbeat watchdog (edge-local holding, ADR 0008) Autonomy: log warning, runtime continues autonomously; if the partition outlasts availability.holdGraceSeconds the embedded SFC engine runs the armed hold chart to a deliberate sequenced safe state and keeps holding (no fence — Autonomy is the legitimate single writer). Failover: runtime self-fences on lease expiry, dated from the acknowledged-renewal age the operator reports in every POST, so neither a symmetric partition nor a reverse-path (asymmetric) one can grant the runtime time past the operator's own clock (#579, #1792, #1802) — running the armed hold chart as a bounded sequence first, then self-fencing so the re-bind margin (≥ hold duration) keeps one writer; operator may re-bind after the safety margin Hold grace then pod-restart dependent (Autonomy); lease duration + margin (Failover). Bench-validated in both modes. Failover (n=3, #942 drill 2b): expiry declared 17.4s, 27.4s and 28.8s after the cable came out, with the re-bind margin exactly 15.0s on every rep. Autonomy (n=1, #942 drill 4): the watchdog fired at +56.8s and the armed hold chart settled the outputs at +67.1s Autonomy: operator cannot issue new commands; the runtime drives the unit's own CM tag space to a sequenced safe state and holds. Failover: bounded hold sequence, then outputs held on self-fence, then failover to a standby. On reconnect the runtime's self-held status reconciles to Held; recovery is an explicit ISA-88 Restart
Controller link fault (cable / NIC / switch port) to a live node Bonded (recommended, ADR 0026): bond driver detects member link-down; no control-plane symptom. Single-homed: indistinguishable from a node partition — watchdog HTTP timeout / lease renewal failure Bonded: traffic moves to the surviving member; the runtime keeps regulating and keeps publishing. Single-homed: escalates down the network-partition row above — hold chart and/or self-fence, then a failover of a controller that never actually failed Bonded: sub-second, below the runtime. Bench-validated (n=3, #942 drill 9). Three reps pulled the member the bond was actually using, and the bond moved to the survivor on every one. No Hold, no lease transition, no self-fence and no re-bind, with dcs_cm_fb_network_running at 1 across 361 one-second scrapes. Two of the three carry the field half as well: the historian's largest hole was 0.104s and 0.148s against a 0.100s cadence, at zero non-Good quality, and the loop's PV held 59.96-60.06% throughout. Single-homed: as for network partition, which the drills reached by other injections and never by pulling a lone cable. See the field-path limit below Bonded: none — no control or data gap. Single-homed: a full Hold/failover cycle for what is only a cable fault. This row is the reason link redundancy is the standard posture
Edge node failure — Autonomy (default) Kubernetes node controller (40s default), measured on the bench at 35-50s; then a hardcoded 30s runtimeCrashGracePeriod that starts when the pod loses PodReady, which on a dead node is the node controller's act and not the kubelet's (#942 drills 1 and 5); unit watchdog/self-Hold Unit lands in Held; failover is a manual operation (dcs unit failover --confirm-fenced once the node is powered off / field-disconnected) re-binding the runtime to a designated standby Manual (operator-paced). Bench-validated (n=3, #942 drill 5). Three power-loss reps measured 87.6s, 96.5s and 104.4s from the cord leaving the PDU outlet to the standby controlling, and the operator's own decision to fail over sits inside those figures. The re-bind after the command was flat at 15.1s, 15.2s and 15.2s, which is the one interval here the product answers for end to end. The spread is all in the term ahead of it: the unit stopped reading Running at 72.4s, 81.1s and 89.2s. Each figure is a lower bound, because the drill's zero is a keypress at the rack and the operator can only be late One unit's control loop down until failover or hardware replacement; running phase survives in Phase.Status. The re-bind itself writes what the coupler is holding on every output whose channel declares a readbackAddress, and nothing at all on one that does not (ADR 0082)
Edge node failure — Failover (opt-in) Lease expiry observed by the operator Runtime self-fenced at lease expiry; a unit still Running mid-batch is placed on ISA-88 Hold by the operator itself once expiry outlasts the hold bound (½ lease, clamped 5–30s); operator auto-re-binds to a free eligible standby after lease + safety margin (½ lease, clamped 10–60s); new epoch stamped Lease duration + margin (tens of seconds). Bench-validated (n=3, #942 drill 1). Three power-loss reps measured 55.6s, 55.9s and 57.6s from the cord leaving the PDU outlet to the standby controlling, inside the predicted 45-60s band. Each figure is a lower bound, because the drill's zero is a keypress at the rack and the operator can only be late Held outputs during the gap, then the unit is re-bound to the standby and the procedure running on it holds with it. Until ADR 0082 the re-bind itself broke that sentence: the promoted runtime's first scan wrote a compile-time default, measured on the bench as zero on two live channels about four seconds before the unit reported itself controlling. It now adopts what the device is holding, and withholds the write where no channel declares where it reads back. Since #1935 it asks the device beside the scan, so the replacement's lease endpoint is up before a coupler it cannot reach has timed out. Since #1936 the drivers' own first dial is beside the start too: one Modbus dial against a coupler the node had no route to was 3s of the 15s to grant on the 2026-09-03 bench failover, per IOModule in sequence, and Start now returns with the dial in flight. Fixed and not yet bench-verified. The batch does not resume on its own: recovery is an explicit ISA-88 Restart
Controller node carrying a network IOModule's io-probe fails IOModule status.state goes Unknown; the AlarmTypeSystem alarm names the probe as what stopped answering Nothing at the equipment level, and that is the response. The Controller a modbus, opcua or ethernetip module names is where its io-probe pod runs. The device itself is reached over the wire (ADR 0021), so the unit runtime keeps scanning it over its own connection. Unknown is outside the auto-HOLD set (ADR 0053 as amended by #1719) Bounded by the io-probe's own rescheduling, which for a controller-bound probe means the node returning No control impact, and no batch is held. What is lost is the DCS's independent view of the device: dcs io read against that module fails while the runtime's own reads keep working. Read the runtime's per-module verdict at GET /api/v1/sites/{site}/units/{name}/runtime/diagnostics. A protocol: simulation module is the exception — its Controller is its device, so it goes Offline and does hold a running batch
Lease loss / self-fence (Failover mode) Runtime local timer against the anchor the operator reports in every renewal — its own acknowledged-renewal age, which is the anchor the Expired decision uses too, so the two expire together (#579, #1802). A reported age that climbs past the lease answers the other fault, a runtime still taking POSTs while the operator gets no usable response back; a sustained reverse-path drop does not reach it, measured on metal (#1792). dcs_runtime_fenced == 1; operator marks lease Expired Runtime drops FB output writes (reads continue), reports NOT_SERVING, publishes retained fenced status; critical Alarm + AuditRecord Immediate (self-fence) Field outputs hold last value / device fail-safe; one-writer guaranteed before any standby takes over
Gateway pod crash Kubernetes liveness probe Pod restart; WebSocket clients must reconnect 10-20s UI temporarily unavailable; no control impact
Control Module FB network down (single CM) dcs_cm_fb_network_running == 0 + retained MQTT health topic Runtime publishes raw driver reads at 1 Hz with quality: "Raw" so CM PVs stay visible; ControlProgram reconciler redeploys on generation/pod change Runtime-dependent PV engineering scaling unavailable for that CM; operator sees raw value flagged Raw; other CMs unaffected
Single FB block error (e.g. transient driver write fault on one address) CM health state Degraded + dcs_cm_program_degraded_total increments on the program's first faulted scan. dcs_cm_blocks_faulted rises and dcs_cm_block_faults_total{block} increments on the block's own edge, so a second block faulting under a standing one is still seen (#1920) Faulting block is skipped and flagged; the scan continues and keeps driving (and interlocking) every healthy output — the program does not halt (ADR 0009) Self-clearing when the block recovers (next scan) Faulted output holds its last value; all other regulation and protection continues; persistent degradation is the operator signal
Control program deliberately removed Operator/reconciler removal Output blocks driven to their configured fail state once before teardown (failState: safeValue default, or holdLast; ADR 0009, IEC 62443 SR 3.6 #972) Immediate Outputs go to their configured predetermined state on removal; a hot-swap redeploy holds instead, to avoid glitching live outputs
Graceful runtime stop for a redeploy (rolling update, image bump, node-pressure eviction) Pod recreated by the reconciler Nothing is written: output continuity is preserved on purpose, the same reasoning the hot-swap path uses (ADR 0009, #1283). The replacement pod restores the commanded values and block operating points it was carrying, before its first scan (ADR 0080) Runtime returns in seconds Outputs hold their last commanded value across the gap. Until #1776 they did not: the outgoing pod held them and the replacement wrote a compile-time default on its first scan, so the field was bumped on every deployment by the path that exists to avoid bumping it
Graceful runtime stop for an outage (terminal stop armed) POST /api/v1/stop/arm on the runtime before the pod is stopped Output blocks driven to their configured fail state once, while the drivers are still connected and before the runtime exits (ADR 0009 amendment, #1283) Bounded by the safe-write budget, inside the pod's termination grace period Outputs reach their configured predetermined state; the runtime queues a safe-stop record that the operator materializes into an AuditRecord, so an outage's trail can assert the field was safed
Ungraceful runtime process death (OOM-kill, node loss, kill -9) No software runs on this path The device's own comm-loss watchdog acts, if one is declared. Declare it with spec.failSafe on the IOModule (ADR 0068); an IOModule that declares nothing keeps whatever its device shipped with, and says so through the FailSafeDeclared condition. The declared timeout has a floor of 45 s, three io-probe read cadences, because any traffic feeds the watchdog and a shorter deadline fires on a routine gap (ADR 0071) Declared timeout What a device can do is per model. A WAGO 750 coupler clears the whole node and offers no substitute value; a cleared 4-20 mA output sits at 4 mA, which reads downstream as a valid 0% and not as a fault

Continuous control and data integrity across a controller failure

The failure table above is written from a batch and procedural perspective, where ISA-88 Hold/Restart is the sanctioned exception path and a bounded gap is immaterial. Continuous control at the control-module level has a different exposure, and it is worth stating plainly.

Most of the layering already protects the record:

  • A broker, historian, or downstream-network fault is covered by store-and-forward. The runtime buffers and replays, so there is no permanent record gap. The one thing store-and-forward cannot cover is a broker that is reachable and refuses the topic. That verdict is the same on every attempt, so holding the message would cost every message behind it. It is dead-lettered to the node's disk and counted on dcs_runtime_mqtt_queue_dead_lettered_total instead (#1780). A non-zero rate there is a record gap, and it is an ACL to fix.
  • A refusal that names the ACL is the one exception to the exception. It comes off the head of the queue the same way, so nothing behind it waits. It is then held in deferred.jsonl and offered again for fifteen minutes before it is given up on (#1814). An upgrade is when an ACL moves, and a bench upgrade discarded 283 messages four and a half minutes before the release that authorised them took effect. Watch dcs_runtime_mqtt_queue_deferred_recovered_total against dcs_runtime_mqtt_queue_dead_lettered_total{reason="not_authorized"}: the first is what an upgrade got back and the second is what its window ran out on.
  • A single CM's FB network going down while the runtime lives leaves the PV visible: the runtime publishes raw driver reads at 1 Hz flagged quality: "Raw".
  • A cable, NIC, or switch-port fault on a controller's own attachment is masked by a bonded interface (ADR 0026), the recommended posture, and three reps on metal say it is masked completely (#942 drill 9). The bond covers the controller side of the link and stops there. Remote I/O is single-homed: a coupler, an ADAM drop or a PLC each has one port on one switch. A dead field switch therefore takes every device behind it at once, and the controller's surviving bond member has nothing left to reach. A validated bond is controller-side redundancy. Field-path redundancy needs a ring topology or dual-homed devices, and ADR 0026 puts that out of scope.

What is not covered is the death of the controller node itself. On failover, a continuous PID loop stops and re-establishes on the standby: integral state is held in the runtime process and persisted to a node-local hostPath, and a hostPath does not follow the pod to another node. The loop restarts from its configured initial conditions. Where it left off is gone with the node. During the RTO window outputs hold last value, go to their device fail-safe (ADR 0009), or are sequenced by an armed SafeStateChart (ADR 0008), and no reader is on the field. The result is a bounded data gap as well as a control gap.

We do not guarantee zero-gap continuous control or an unbroken record across a controller death. For most processes this is immaterial. For a critical continuous parameter under 21 CFR Part 11 / ALCOA+ it may not be. What the product does guarantee is that the gap is bounded, timestamped, and cause-attributed, bracketed by a Hold event, a failover AuditRecord, and an explicit ISA-88 Restart. The bracket is what makes the gap a documented one. It is also carried inside the batch production record. At batch-terminal time the reconciler materializes each failover into the record's spec.failoverEvents[], with the control gap (Hold → Restart) and the data gap (last acknowledged contact → lease re-established) bounded separately and each bound referencing its evidencing AuditRecord. A QA reviewer reads the explained gap in the record itself, with no operator log to chase (ADR 0041, #1308, and Batch Production Records).

The two gaps are not bounded to the same tightness, and each gap says which it is (#1805). The control gap's bounds are its own edges. The data gap's contain it wherever it opens at an acknowledged contact: the hole in the historian lies inside the published interval and data may be present near either edge. It opens at the last lease renewal or heartbeat the control plane had acknowledged. That is the last instant the control plane can evidence the runtime was alive. It closes at the lease re-establishment on the replacement. By the one-writer guarantee the replacement reaches that point after it has already resumed publishing. Opening it at the lease EXPIRY instead published a bracket offset from the hole it explains: 28.0 s late at the opening and 8.01 s long at the closing, against a 45.2 s hole, measured on drill 10 below.

A failover the control lease did not drive produces neither lease record. Neither of the two paths that reach it is exotic. A planned-maintenance failover is taken while the old lease is still held, and a unit in Autonomy mode holds no lease at all. There the closing bound is the re-bound runtime returning to normal, and the opening is the same last acknowledged contact, carried on the re-bind record (#1753, #1805). Where nothing was ever acknowledged the opening falls back to the record of the loss itself. That bound is late, so the gap's bounds field reads Overlapping and the absence reaches back before the published start (#1822). The record discloses which fallback was taken on its PartialData condition. It publishes no start time it cannot stand behind, and no bound kind it cannot stand behind either.

If your process needs the unbroken record, the decided answer (ADR 0041) is redundant-collector sourcing: a redundant PLC or OPC UA server with historical buffering that the historian backfills on reconnect. That keeps the record whole even when control briefly holds, and in regulated contexts the record is more often the binding requirement than bumpless regulation. Raise it during design. It is a deployment topology, and no product setting turns it on. Hot-standby FB-state replication is deferred behind the reversal triggers ADR 0041 records.

What the RTO figures measure

The figures in the table measure the equipment: lease expiry, fencing, re-bind, and a runtime driving the field again. When the batch starts moving is a separate number. The two are deliberately different.

A failover holds the unit, and since #1723 that hold reaches the procedure running on it. The phase, its operation, its unit procedure and the batch all reach Held, and they stay there. Nothing in the product resumes them. Recovery is the explicit ISA-88 Restart this page has always specified, issued by someone who has decided the equipment is fit to run. The wall-clock time to that decision is a property of your staffing, and the control system puts no bound on it.

Plan for it. A unit that fails over unattended stays held until an operator arrives. A dwell interrupted by the failover resumes with the held interval added back, so it keeps the time it still owes even though nobody was driving. That is the correct posture for equipment taken off closed-loop control, and it does mean an unattended run now stops where it once rode through.

What the bench has measured

The #942 drill campaign exercised these failure modes on real hardware and is complete. Every row it set out to reach carries its rep count in the Recovery Time column above. The measured-results table, drill by drill and including the reps that were thrown out, lives in cndcs-deploy-bench drills/ha-drill-runbook.md § Results, and #942 carries the campaign record under the bench sprint #400. The findings below are the ones that change what this page claims.

Detection is a renewal-phase offset. A Failover lease expires 30s after the last successful renewal, and renewals run every 5s. They ran every 10s until #1909, which is the cadence the reps below were measured at. Detection measured from the fault is therefore 30s minus however old the lease already was when the fault landed, and where in the renewal cycle a fault lands is nothing a deployment controls. Three cluster-path partition reps measured 17.4s, 27.4s and 28.8s, and that 10s spread is one renewal interval at the cadence of the day. The system constants are the 30s lease and the 15s re-bind margin. The margin was exact on all three reps, one of them logging expiredSince 15.009s against a holdBound of 15s. What a deployment sees on top of those constants is a 0-10s offset it cannot tune, so publish detection as a range. The true zero is still not observable, because the fault is not an event the control plane sees. The last successful renewal now is, and it is the tightest bound on the fault that exists here. Since #1805 the operator stamps it into the LeaseExpired audit record, where the batch production record reads it to open the data gap. Unit.status.runtimeBinding carries only lastTransitionTime, which is the last state change and not a renewal.

The runtime's anchor is one renewal older than the operator's, and the renewal cadence has to be measured against it. Every renewal POST carries the operator's acknowledged-renewal age. That age is measured before the attempt goes out, so it names the renewal before this one. The runtime dates its lease from it. A Failover runtime therefore self-fences one renewal interval before the operator declares the lease Expired, and that lag is deliberate. The runtime cannot know its own response reached the operator, and crediting itself for a 200 that may have been dropped is what #1802 removed. The direction is the fail-safe one. The one-writer inequality above is unaffected, because what the lag does is fence the runtime earlier.

The consequence is that the deadline the renewal loop races is the runtime's and not its own. Until #1810 the loop computed what was left of the lease from its own last success, which is one interval younger than the anchor it had established. The two errors cancelled exactly, so the clamp #1802 added returned the interval it was handed and never once lowered it. A single missed renewal then put the retry precisely on the runtime's deadline at every lease duration from six seconds up, and half a second the wrong side of it at the five-second minimum. Measured at five seconds, one refused renewal fenced a healthy unit twice and ran the ADR 0008 safe-state hold on the plant both times, for a second of dropped output writes on a plant nothing was wrong with. The loop now halves what is left of the runtime's own anchor, so a retry lands inside it until the lease is genuinely spent. The steady cadence is unchanged at every duration of six seconds or more, including the bench's 30s lease.

The tail of a failover RTO is the replacement pod becoming Ready. The operator declares recovery on a conjunction: the lease re-established on the new node, and the runtime pod reporting Ready. A conjunction is satisfied by its later term. On four measured automatic re-establishments the lease landed first every time, by 2.7s to 3.4s (#1752). Those margins are upper bounds, because the log line marks the reconcile that observed readiness, one step after the readiness edge itself. What the reps settle is the ordering, and it means pod start gates the published figure. The lease machinery is already finished while the RTO is still running.

Two of the three terms in a manual failover's detection are not the product's. The Autonomy row's detection is a chain of three, and only the middle one is a constant. Kubernetes has to notice that the node stopped renewing its own node lease. That lag measured 49s and 50s on the two drill-5 reps where the node events were recoverable, against 35s to 42s on drill 1. Only then does the pod lose PodReady, which is what starts the hardcoded 30s runtimeCrashGracePeriod. The operator's own reconcile then added a further 2s to 10s. The three terms sum to the 72.4s, 81.1s and 89.2s measured before the unit stopped reading Running (#942 drill 5, n=3). A correction drafted against this row in August 2026 held that the node controller did not govern and that the 30s grace decided the figure by itself. It was drafted against reps that had measured a graceful shutdown, where the kubelet terminates the pod at T0 and the lag term really is zero. On a power loss the lag is the largest of the three, and the published row was right to name it.

A sequenced safe state waited on an audit publish. The one Autonomy partition rep taken so far ran the unit's own safeStateChart and settled the outputs 10.3s after the watchdog fired, against a three-step chart that should take roughly 300ms at a 100ms scan. The hold path emitted its triggered audit event before calling the SFC engine, and that emit published at QoS 1 on a context with no timeout, against a broker the partition had just taken away. The stall sat ahead of the engine, so neither WRITE had run. The field held its pre-fault values for the whole 10.3s. #1779 moved the emission onto the store-and-forward queue, which makes it a disk write. The rep above was taken on a release that predates it, so the corrected figure waits on drill 4's two remaining reps (#1791).

The edge safe state was a one-shot per runtime process. Drill 4's two remaining reps were taken back to back on one runtime process on 2026-08-25. Rep 1 did everything ADR 0008 promises: the watchdog fired at T0 plus 58.1s, ran the unit's safeStateChart, and settled 0.3s later. Rep 2 ran nothing at all. There was no watchdog line, no hold, no hold event on the edge's disk and no AuditRecord, and both outputs read HELD_LAST_VALUE for the whole partition. The latch that makes the watchdog fire once per partition is cleared by POST /api/v1/hold/release, and the phase reconciler used to send that only when Status.EdgeSelfHeld was set. That flag is written from an HTTP GET to the runtime pod, and a partition puts that pod out of reach. The recovery Restart therefore sent nothing. Every instrument reported the unit as covered throughout, because each of them is about the armed program and not about the watchdog behind it. Two reps on one process is the only way to see it, which is why three earlier drill-4 attempts and the 2026-08-23 valid rep never did. #1818 puts the release on any command out of Held and adds the signal that separates armed from able-to-fire.

And nothing an operator opens showed any of it until #1833. Both amendments above put the arming on Unit.status.runtimeBinding.edgeHold. It reached no gateway page and no CLI command, so confirming that a unit was covered meant kubectl get unit -o jsonpath. Three surfaces answer it now.

The Unit detail page carries an Edge Hold row, and dcs get units carries an EDGE HOLD column. Both read that record. Both therefore keep answering while the pod is unreachable. dcs get runtime prints the runtime's own live arming as Armed Program, Baseline Slot and Edge Watchdog. That reading is fresher, and it stops with the pod.

That was not true of the page until #2149. The record kept answering and the page did not paint it. The Unit page re-reads every ten seconds, and one of its reads is the runtime's own diagnostics. A build missing that read was dropped whole, on the rule that a tick must not swap a degraded page over a healthy one. A partition is the case where that read always fails. On the bench on 2026-09-21 a partitioned Failover unit read Running, Ready and Held on ipc-2 for the 36 seconds the partition lasted, beside reported 6s ago on the Edge Hold row. The Unit record carrying Expired had reached the browser and was discarded. The lease expiry and the 15 seconds before the re-bind were never on screen.

The Runtime Diagnostics section now degrades in place. It reads Not read — the runtime did not answer on a grey dot, because nobody could ask, and a Last Answer row dates the silence from the last answer that page had. The rows read from the control plane keep moving: the state badge, Runtime Ready, Availability and Edge Hold. The repaint trails the record by up to five seconds during the outage, because the gateway waits that long for the runtime before it answers 502.

Six readings are separated, and two of the six are byte-identical on the wire. nothing armed on a unit that declares a safeStateChart is a chart that failed to arm. The same bytes on a unit that declares none are the configured posture. not reported says a report has never arrived, which is an instrument to repair. armed, cannot fire is the latched watchdog this drill lost a rep to.

The record is the promoted runtime's own, or it is not reported. A manual failover of an Autonomy unit on the bench left the row at nothing armed for as long as the heartbeat loop lived after the re-bind (#2137). The loop's memory of the previous runtime's answer had been copied into the rebuilt binding as the new runtime's, and the loop's memory of having armed the previous pod kept the arm from being posted to the new one. A re-target onto a different address now drops both. The row reads not reported until the standby answers. The loop arms and beats at the re-target, and again when the pod turns Ready, instead of waiting for its next tick. A moved report nudges the reconciler, so the row follows the standby's answer by seconds. An Idle unit has no requeue of its own, and before this the row waited on whatever next happened to reconcile it.

The flag answered to a Restart that was never coming. Drill 12 rep c on 2026-09-02, chart 0.7.1, fenced bench-loop-01 for 9.5 s (#1909). The runtime ran its bounded hold, dcs_runtime_self_held went to 1 at 20:56:28.7, the lease came back and the fence cleared at 20:56:38, and the gauge stayed at 1. The batch that was fenced had completed before the cord came out, so the hold ran on an idle unit under the baseline. No phase was Held by it, and no Restart was coming. Two batches ran to Complete on that runtime and it still read 1 at 22:40. Rep 2's row carried self_held=min=1 max=1 through a rep in which nothing held. #1804 had given a hold the pod found a road back, and left a hold the pod ran with the Restart alone. #1919 adds the two roads the work itself takes: a phase starting on the unit under work that is not its own, and the engaged phase ending. Both release the watchdog latch too, because in Autonomy the same shape leaves the next batch with no edge cover at all.

A bonded member pull is invisible above the driver, and the positive control is what makes that worth believing. Drill 9 asserts that nothing happened. A rep that pulled a dead cable, pulled the standby member or reseated the plug before the driver committed the loss reports every one of the eight observables clean. Each rep therefore reads the bond's own Link Failure Count on both sides of the pull and requires the named member's count to have risen. It read a delta of exactly one on all three reps, on the member the cue named, with no second member moving. Behind that control the control plane saw nothing at all: no Hold, no lease transition, no self-fence, no re-bind, and dcs_cm_fb_network_running at 1 on every one-second scrape. The field saw nothing either on the two reps where it is readable, with the historian's largest hole at 0.104s and 0.148s against a 0.100s cadence and no sample below Good. None of it says anything about the field path. The bond is on the controller, the coupler has one port, and the switch those devices hang off is common-mode for every one of them at once.

A reverse-path partition does not reach the gate built for it, and the one POST that does land used to make things worse. The reverse-path case is the asymmetric one. The operator's renewal POSTs still arrive and the runtime's responses never come back. #579 closed it by having the operator report its acknowledged-renewal age in every POST. The runtime then self-fences on that report, and no longer on bare POST receipt.

The bench reps (#942 drill 3, 2026-08-24, n=3) say that gate is not reached at all. TCP needs the return path for its ACKs, so with none coming back the operator's send window fills and the kernel puts no new data on the wire. On rep 1 exactly one renewal POST landed after the cut, 2.7s in, and none after it. All three fenced on the runtime's own local timer through the ADR 0008 pre-fence hold. That is the symmetric partition's path, and the operator's lost-return-path report appeared on none of the three.

The transport decided the fence timing, because that one POST arrived carrying a fresh acknowledged age and the runtime dated its lease from the arrival. With mTLS on the channel is HTTP/2, so the POST rides a connection that already exists and lands, and it bought the runtime up to a renewal interval past the operator's own clock. Measured, the runtime fenced 10.7s AFTER the operator had declared the lease Expired, and 4.4s before the re-bind decision, where the design intends the runtime to have fenced first with the whole 15s margin behind it. Nothing dual-wrote, because the standby took the lease 14.6s after the fence. With mTLS off the channel was HTTP/1.1 and the broker closed each response body without draining it. Go's transport returns a connection to the idle pool only once the body reader has reached EOF, so every one of those was retired on the spot. The renewal then needed a handshake the cut return path cannot complete, so no POST landed and the two clocks stayed aligned by accident (#1801).

Since #1802 the runtime dates its lease from the age a report carries instead of from the arrival of the POST that carried it, so a landed POST extends nothing and both postures fence at the same instant. That ordering is what made #1803 safe to take. #1803 drains those response bodies. The plaintext connection is pooled, the no-TLS posture lands its one POST as well, and the accident above is gone. The two postures now differ in how the renewal travels and in nothing else. On the pre-#1802 runtime the same drain would have handed the no-TLS posture h2's late fence, as a side effect of what reads as an efficiency fix.

The one-writer guarantee now rests on one inequality, and on no renewal interval or POST timeout at all. The re-bind margin is at least the hold bound. The hold is budgeted from the lease deadline, so a late notice spends the runtime's own safe-state time and never the standby's. Before that the guarantee rested on the arithmetic difference between two unrelated cadences: 4.4s on the bench's own rep, and 3.3s at the tightest point of the configurable range. It moved whenever either cadence moved. The fix ships in v0.6.0, and the bench has run it since v0.6.1 landed on 2026-08-25.

Re-exercised on metal on chart 0.6.3, 2026-08-27 (#1824), three reps. The ordering is no longer inverted. The runtime's self-fenced line still names trigger=local timer on every rep, so the #579 ack-age gate is still never reached on this partition. The fence now lands 3.2s, 3.4s and 3.4s before the operator declares Expired, reversed from the pre-fix reps' +10.7s and +10.9s after. The re-bind margin read 18.2s, 17.9s and 17.5s, above the 30s lease's 15s hold bound on every rep, where the pre-fix reps left 4.4s and 4.2s. cndcs-deploy-bench drills/ha-drill-runbook.md § Results drill 3 carries both sets.

RTO figures for the rows no drill reached are still derived from lease and probe timings and exercised on kind and in CI. Every row this section once listed as waiting on a drill of its own has been measured on metal, both drills that stood short of the campaign's three-rep bar reached it, and the campaign closed with every drill taken.

The continuous loop was the last of those drills to be taken (#1794, #942 drill 10, three reps on 2026-08-24 and three more on 2026-08-25). The control gap and the data gap came out as different quantities, and the difference has a mechanism behind it. The data gap is the hole in the historian, and across the six reps it ran 43.0 s to 52.8 s over 214 to 263 missing samples at a 0.200 s nominal cadence. The control gap is the interval in which nothing was writing the field, and it is the longer of the two on every rep: by 8.61 s on the first sitting and by 7.32 s on the second. That interval is the one-writer guarantee working. The replacement runtime is up and publishing readings before it holds the lease, and it does not drive the field until it has one. A single combined RTO would have hidden the distinction.

That separation is quoted as a difference on purpose, and it is only a difference when both gaps are anchored on the same instant. The data gap is bracketed by two published samples. A control gap measured from the moment the cord came out carries however long the walk back from the rack took. On 2026-08-25 one rep's key landed 11.8 s after its own last sample, which is enough to reverse the ordering when the two numbers are set side by side. The two sittings also closed the control gap with different witnesses: the first watched the field move, the second read the promoted runtime's own first write (#1794). A field-side witness cannot close earlier than a runtime-side one, which is the direction the two separations differ in.

The re-establishment bump has a figure now. It is the architecture's number and not a defect's. Measured on the three reps of 2026-08-25, on a runtime carrying both ADR 0082 and #1815: peak deviations of 0.0135, 0.0096 and 0.0116 against a band sigma of 0.0251, 0.0247 and 0.0259, each back inside the band within one 0.200 s sample. That is a re-establishment and not a step response, because the loop was still on its operating point when control resumed. Its setpoint was re-commanded to 60.0000 while the PV already read 59.9872. On the pin before it the same measurement was 53.95 % against a plant held at 59.95 %, which is the defect #1815 fixed and not the architectural gap. Taking the figure on a runtime that predates either fix would have filed a repaired defect against the architecture.

The same three reps also put the published bracket beside the hole it explains. With #1805's fix in, the interval the record carries in spec.failoverEvents[].dataGap contained the whole hole on all three, at 112 %, 135 % and 131 % of its length, opening 1.4 s to 7.8 s early and closing 4.8 s to 7.2 s late. Both bounds err outward, which is what bounds: Containing claims to a reviewer, and it is the claim the paragraphs above make on the product's behalf.

One figure on this page has not caught up with its drill. The Autonomy partition beyond holdGraceSeconds reached its third valid rep on 2026-08-25 (#1791), and the watchdog and hold-chart timings quoted in the network-partition row above are still the single reading taken before them.

IEC 62443 Availability Targets (FR 7)

Component Availability Target Basis
Control operators 99.9% (8.7h/year downtime) Leader election + 2 replicas. Neither is a full answer to a loss of the Kubernetes API, which reaches every replica pinned to the lost apiserver endpoint, so the physical-operator no longer exits when its manager stops. It stands down, keeping edge liveness granted for physicalOperator.leaseStandDown (#1894)
Unit runtime (Autonomy) 99.5% (43.8h/year downtime) Single pod per unit; restart + replay; manual fenced failover on node loss
Unit runtime (Failover) 99.9% (8.7h/year downtime) Adds lease-based self-fence + automatic re-bind to a designated standby (ADR 0006); requires standby hardware with field-network reach
Gateway 99.9% Multiple replicas + autoscaling
MQTT broker 99.5% (single) / 99.9% (HA) Single instance by default; optional HA mode (mqtt.ha.enabled) deploys StatefulSet with configurable replicas and per-instance storage

These targets assume the HA profile (values-ha.yaml) is deployed.

Deploying for High Availability

Enable HA by applying the overlay values file:

helm install dcs ./deploy/helm/cloud-native-dcs \
  -f deploy/helm/cloud-native-dcs/values-ha.yaml

A replica count answers the loss of a node and only sometimes the loss of the Kubernetes API. Both physical-operator replicas reach the apiserver through the same Service address. Each pod's connection is pinned to one apiserver endpoint, so the follower rides an outage through when it is pinned elsewhere and dies with the leader when it is not. Acquiring the leader lease is itself an API write. What covers the second case is physicalOperator.leaseStandDown. It keeps a leader granting edge liveness for 90s after its manager stops (#1894), and it is on by default, in the HA overlay and outside it.

The standby has the same pin and nothing to stand down from. When the node it is pinned to loses power, nothing closes its connection. client-go keeps sending the election's reads down it until its HTTP/2 health check gives up, which by default is 45s. That is longer than a unit's whole lease. physicalOperator.apiserverHealthCheck sets the check to 2s and 2s on this operator, so the standby dials a live apiserver inside 4s and the election runs from there (#1930).

This enables: - 2 replicas for every operator - PodDisruptionBudgets (auto-enabled when replicas > 1, with explicit enable also supported) - Gateway autoscaling (2-5 replicas) - Pod anti-affinity (spread across nodes) - PrometheusRule alerting - MQTT HA mode (StatefulSet with per-replica storage, enabled with mqtt.ha.enabled=true)

Alert Runbooks

DCSRuntimeUnhealthy

Severity: Critical Condition: dcs_runtime_healthy == 0 for > 2 minutes

What happened: At least one field I/O driver on a unit runtime is disconnected, or the operator could not read the runtime's status at all.

The gauge reads 0 for a partial loss as well as a total one (ADR 0078). A partial loss Holds the batch, so a gauge that stayed at 1 through it would be an alert reading all-clear over a plant that had stopped. The runtime's own gRPC readiness verdict is narrower: it reports NOT_SERVING only after every field driver has been gone for the grace period. Its liveness verdict is a separate service and never fails on a field loss, because restarting the pod does not reconnect a cable (#1918, ADR 0006 2026-09-02 amendment). It fails only for a network that says Running and has stopped scanning.

Simulation drivers are not counted towards this (ADR 0075). A fully simulated unit reads healthy because it has no field connection to lose, so this alert firing always means either a real endpoint has gone or the status API stopped answering.

To tell the two extents apart, read dcs_runtime_driver_connected for the unit. It carries a series per driver.

The same count is on every surface a person opens. Three of them read 2 of 3 connected off the runtime's own field-driver census: the Field I/O row in the Diagnose panel's Runtime tab, the row of that name on the Unit detail page, and the Field I/O line of dcs get runtime. That census is what this alert's verdict is taken over (#1844). Until then those rows carried the runtime's primary driver, which is a simulation driver on every deployment. All three answered simulation and Connected: Yes whatever the plant was doing.

Actions: 1. Check the Diagnostics page for runtime status 2. Check driver reconnection metrics: dcs_driver_reconnect_failures_total 3. Verify field device connectivity (network, power, protocol endpoint) 4. If the runtime pod was restarted by kubelet, check its startup log for one restored network line per program the unit was running. no persisted network found is not the line to look for: the singleton replay path that logged it was deleted in #1748

DCSRuntimeDriverDown

Severity: Warning Condition: dcs_runtime_driver_connected == 0 for > 5 minutes

What happened: A specific I/O driver has been disconnected for an extended period. The ReconnectingDriver is attempting automatic reconnection.

Actions: 1. Check which driver: look at driver_name and protocol labels 2. Check IOModule status: dcs get iomodules -s <site> 3. Verify the field device is reachable on the network 4. Check for protocol-specific issues (Modbus TCP port 502, OPC UA port 4840)

DCSRuntimeDriverConfigRefused

Severity: Warning Condition: dcs_runtime_driver_config_valid == 0 for > 1 minute

What happened: The runtime accepted the driver and refused its configuration. This is not a connectivity failure and it does not look like one. The driver connects, answers every read, and reports dcs_runtime_driver_connected == 1. What it does not do is run the profile the module declares, so a simulation module holds every channel at its initial value and injects no faults. The plant presents as a set of tags that have stopped moving.

The reachable causes are a behaviour or fault parameter the driver would not parse, a dependency cycle among behaviours, and a profile whose JSON does not unmarshal. The gateway refuses each of those in an authored document (ADR 0064). A module in this state usually arrived through kubectl apply, through POST /api/v1/apply, or from a document written before that guard shipped.

Actions:

  1. Read the message. It is not a metric label. dcs get runtime -s <site> --unit <unit> prints it in the CONFIG column, and the Diagnose panel's Runtime tab shows it in the I/O Drivers table.
  2. Fix the IOModule's spec.simulation block or the SimulationPreset it binds, and apply it. The unit controller pushes the corrected config to the running runtime, which re-applies the profile without a pod restart.
  3. Confirm the gauge returns to 1 on the next watchdog poll.

One failed behaviour costs the whole module's profile, so the message names the first refusal and the module runs none of its behaviours until that is fixed.

DCSOperatorDown

Severity: Critical Condition: One or more operator up metrics absent for > 5 minutes

Actions: 1. Check the Diagnostics page for operator health status 2. Contact your system administrator if operators are unhealthy 3. Verify RBAC: missing permissions cause silent failures

DCSReconcileErrorRate

Severity: Warning Condition: Reconciliation error rate > 0.1/s for 5 minutes

Actions: 1. Check operator logs for error messages 2. Common causes: missing CRDs, RBAC issues, etcd timeouts 3. Check if a recent deployment introduced a regression

DCSMQTTQueueBacklog

Severity: Warning Condition: MQTT queue depth > 1000 messages for > 5 minutes

Actions: 1. Check the Diagnostics page for MQTT broker status 2. Verify broker connectivity from runtime: check dcs_runtime_mqtt_publishes_total, whose outcome label separates what reached the broker from what is queued or dropped 3. If queue is approaching the 10K limit, messages will be evicted (check dcs_runtime_mqtt_queue_evicted_total)

DCSWatchdogHoldTriggered

Severity: Critical Condition: Watchdog Hold triggered in the last 5 minutes

What happened: The unit controller's watchdog found field I/O missing while a batch was actively running, and placed the unit on ISA-88 Hold to prevent uncontrolled process behavior.

The alarm says which extent. A watchdog-<unit>- alarm at Critical means every field driver was gone. An io-loss-<unit>- alarm at High means some were still answering, and it names the ones that were not (ADR 0078). Both Hold, because a batch writing to the channels that still answer and failing silently on the rest is the worse position of the two.

The message follows the outage. Another bus going, or one coming back while the loss is still partial, rewrites the message on the alarm that already stands. The drivers it names are the ones that are missing now. The alarm keeps its acknowledgement and its creation time through that, so the interval from the raise to the clear still measures the one outage.

The Hold stands until you lift it. The unit carries status.conditions[HeldByRuntimeFault], which records that the physical layer took the equipment off closed-loop control and bars the procedural layer's level-triggered drift repair from reading the Held state as a lost command and restarting the unit itself. The drivers coming back clears the alarm and does not resume the batch: nothing but an operator has confirmed the equipment is fit to run. The Restart you issue answers that marker. The phase records which one it reflected, so it does not re-read the same marker as a fresh fault while the procedural operator's view of the unit is still catching up. The UnitProcedure records when it left Held, so a unit the Restart never reached is restarted again (#2155).

Actions: 1. Investigate why the drivers disconnected (see DCSRuntimeDriverDown runbook). The High alarm's message names them 2. Once drivers are restored, issue a Restart command to the held batch running on the unit: dcs command Batch <batch-name> Restart -s <site> 3. Verify the batch resumes correctly 4. Review the Alarm CR created by the watchdog for timestamps and details

The alarm follows the drivers. The controller keeps reading driver health while the unit is Held, so the alarm moves to a Cleared state and stamps status.clearedAt as soon as the drivers reconnect. It also follows the extent: a partial loss that becomes total clears the High alarm and raises the Critical, and a total loss that recedes to partial does the reverse. It does that whether you Restart, Stop or Abort, and whether or not anyone has acknowledged it. A watchdog alarm still sitting Active therefore means the drivers are still down, and the interval between the alarm's metadata.creationTimestamp and its status.clearedAt is how long they were. The RuntimeHealthDegraded condition on the Unit leaves on the same pass.

Unit runtime fenced (control lease lost)

Severity: Critical Condition: dcs_runtime_fenced == 1 for > 1 minute (availability mode Failover)

What happened: A Failover-mode unit runtime could not renew its control lease and self-fenced. It stopped writing FB outputs (reads continue) and reports NOT_SERVING. The physical operator lost contact with the runtime. The unit's RuntimeBinding.leaseState shows Expired and a critical lease-expired Alarm was raised. Field outputs hold last value (or device fail-safe). This is the designed one-writer guarantee at work, and no fault in itself. It does mean that unit is not controlling.

Actions: 1. A brief 1 at runtime startup is normal (Failover runtimes start fenced until their first lease grant). Only sustained fencing is an incident. 1a. Read the alarm's own sentence first. It says which of the two expiries this is (#1890). "Control lease expired: unit X runtime on N self-fenced" is a runtime that held the lease and stopped answering. It is fencing itself on its own clock, and a standby takes over from a writer that has stood down. "Control lease never established: unit X runtime on N did not answer a renewal within of being bound" is a runtime that was bound and never took the lease at all. Nothing was driving that unit and nobody fenced. The fault is in the start-up. The alarm also names what the last renewal attempt got back. A refusal came from the runtime's own handler, so that runtime is up and answering. A transport error means nothing was reached. 1b. Check whether the edge was in the blast radius at all (#1894). A control-plane event that takes the Kubernetes API away reaches every operator pod wherever it sits, and the grantor is one of them. A unit on an untouched node can therefore fence on a platform fault it has no relationship to. The signature is leader election lost in the operator log at the same minute, on a node that stayed up. The operator no longer exits there. It stands down, so a lease that expires anyway means the API was away for longer than physicalOperator.leaseStandDown. Raise that value, and leave the unit's lease duration alone. The lease duration is also how long a genuinely dead controller keeps the plant waiting. 2. Determine whether the edge node is dead/partitioned or the operator↔runtime path is broken: dcs get unit <name> -s <site> (RuntimeBinding + FailoverRequest conditions), node status, and dcs_runtime_lease_expirations_total. 3. Confirm the outputs really are held. dcs_runtime_writes_total{outcome="fenced"} counts the writes the fence dropped, and outcome="ok" counts only the ones that reached a device. The ok series for that unit is therefore flat for the length of the episode. An ok series that is still moving means something is reaching that field while the runtime is fenced. That is the one-writer guarantee being broken, and it is an incident of its own.

The fenced series starts moving one renewal interval before the operator declares the lease Expired. The runtime dates its lease from an anchor one renewal older than the operator's, and that lag is deliberate and fail-safe (ADR 0006). The earlier start is that lag. It is not a clock disagreement between the two sides. 4. If the node is genuinely down, the operator auto-re-binds to a free eligible standby after the safety margin. Confirm a standby exists (dcs unit failover <name> lists eligible targets) and is enrolled (dcs.io/site label) and Ready. 5. If no free standby exists, a NoFreeTarget condition is set. Provision or free a standby, or run a manual co-located failover with dcs unit failover <name> --to-node <node> --allow-colocation --reason "<justification>". --reason is required with --to-node and lands in the audit trail. 5a. If dcs get unit <name> -s <site> reports autoFailover.suspended, with a Critical alarm naming every standby that was tried, the automatic path has re-bound this unit to each of them and none took the control lease (#1890). Re-binding again would repeat an experiment whose result is in hand, so it has stopped. The pod on the current binding keeps trying. The first renewal it acknowledges clears the condition, the alarm and the ledger with no operator action. A suspension is not by itself a reason to intervene. What it is a reason to look at is why no replacement is coming up: a registry the edge cannot reach, an image that will not pull, a driver endpoint that is down, a message broker the replacement cannot reach (its start blocks on the MQTT connect, #1933), or a start-up genuinely slower than availability.runtimeStartupGraceSeconds. kubectl describe pod <unit>-runtime -n site-<site> answers the first three. A container that is Running and never Ready while the broker is down is the fourth. A deliberate dcs unit failover is available throughout and resets the budget. 6. Recovery on the standby is an ISA-88 Restart. Verify the batch resumes and the lease returns to Held. 7. A lease back at Held is permission to write, and not evidence that anything reached the plant. An output whose channel declares no readbackAddress withholds its writes until something commands it (ADR 0082). A runtime can therefore be unfenced, Ready and driving nothing. dcs_runtime_writes_resumed_total moves once per fence episode, on the first output write that reached the device, and the runtime logs the same edge naming the address. dcs_runtime_outputs_unestablished says how many outputs are still withholding, and the Restart in step 6 is what commands them. Read dcs_runtime_outputs_holding beside it (#1816): an output that adopted the coupler's value re-writes that value every scan while ignoring its own input, so it moves the write counters and the resumed edge exactly as a driving one does, and the loop behind it is not regulating. The Restart answers that count too.

The failover raises a high-severity Alarm of its own, alongside the critical lease-expired one, and both follow the runtime. The failover Alarm says the runtime was re-bound and is not back to normal yet, so it clears once the replacement pod is Ready and holding its control lease on the new node (#1733). It does that whether or not anybody has acknowledged it. It does not wait for the ISA-88 Restart in step 5, which belongs to the operator. A failover Alarm still sitting Active therefore means the runtime has not come back. That is a live problem to work. Its metadata.creationTimestamp to status.clearedAt interval ends at whichever half of that condition lands last. In Failover mode that is always pod readiness. The runtime starts fenced and stays fenced until the first lease renewal reaches it. A fenced runtime reports its gRPC readiness as NOT_SERVING, so its readiness probe cannot pass before the lease grant has landed. Its liveness probe asks a different service and passes throughout, so a fence of any length is no longer answered with a restart (#1918). Read the interval as the time to a Ready pod on the standby. It does not apportion that time between pod start and the lease. The lease-expired Alarm opens earlier, at lease expiry. Its span adds detection, the safety margin and the re-bind in front of that. Failover Alarms inherited from a version before 0.5.1 are the one exception to all of this, and the Failover Runbook says how to read them.

DCSCertExpiryWarning

Severity: Warning Condition: A TLS certificate file observed by a DCS component expires in <48h

What happened: cert-manager normally renews certs 8h before NotAfter (24h duration, 8h renewBefore). If the gauge dcs_cert_expiry_seconds - time() has fallen below 48h, renewal has not yet landed, most commonly because cert-manager is wedged or the component is an off-cluster edge node that cert-manager doesn't manage.

Actions: 1. Identify the affected cert: $labels.component, $labels.host, $labels.file. 2. Check cert-manager: kubectl get certificate -A (look at Ready + Renewal). 3. Describe the Certificate: kubectl describe certificate <fullname>-<component>-mtls. 4. Check cert-manager logs: kubectl -n cert-manager logs deploy/cert-manager. 5. Confirm the reloader picked up the last rotation. Grep component logs for TLS certificate material reloaded. 6. For edge nodes (unit-runtime devices joined at the deployment layer), the in-cluster cert-manager does NOT renew. An automated edge renewal loop is not yet implemented, so rotate manually per the rotation runbook.

DCSCertExpiryCritical

Severity: Critical Condition: A TLS certificate file expires in <8h

What happened: Renewal has failed. mTLS handshakes will break at NotAfter. Unit-runtime ↔ gateway, gateway ↔ historian, and MQTT will all cascade.

Actions: 1. Force-renew: kubectl cert-manager renew <fullname>-<component>-mtls. 2. If cert-manager itself is broken, the emergency mitigation is a redeploy with --set mtls.enabled=false. Accept the IEC 62443-3-3 SR 4.1 finding and file an incident. 3. Automatic renewal should have fired at two-thirds of the certificate lifetime. Check the cert-manager controller logs for why it did not, or the forced renewal will expire the same way.

DCSHistorianPruneJobFailing

Severity: Critical Condition: No successful run of the historian prune CronJob in over 12h, measured off the CronJob's own last success, its creation when it has never succeeded, or any labelled Job's completion. A suspended CronJob is exempt (opt-in rule, default off, enabled via monitoring.prometheusRule.historianDisk.enabled: true). Before 0.7.6 the rule also required a failed Job on record, and a Job whose pod never scheduled never fails, so the rule could not see that shape (issue #1942).

What happened: The historian prune CronJob runs every 6h. Two missed cycles let TimescaleDB chunks accumulate beyond the 6h target, and on local-path-style PVs the historian PVC begins to outgrow the node root disk. Left unattended this triggers a DiskPressure cascade across the cluster (observed end-to-end in a 2026-04-22 demo-instance incident, with the prevention guidance codified in Historian Disk Pressure). Two shapes produce the silence. A run that ran and lost leaves a failed Job with logs. A run whose pod the scheduler refused leaves an active Job and a Pending pod, and until 0.7.6 that one Job suppressed every later run.

Actions: 1. Read the Historian Prune card on the System Health page, or dcs health. It names the unfinished run when there is one. 2. Inspect recent Jobs and their pods: kubectl -n {{ $labels.namespace }} get jobs,pods -l app.kubernetes.io/component=historian-prune --sort-by=.metadata.creationTimestamp. A Pending pod with a FailedScheduling event is a placement fault. Name the taint in historian.tolerations. 3. Get the most recent failure's logs: kubectl -n {{ $labels.namespace }} logs job/<name>. Common causes: CNPG primary unavailable, role/grants missing on the prune SQL user, statement timeout under load. 4. Once the underlying cause is fixed, kick a manual prune: kubectl -n {{ $labels.namespace }} create job --from=cronjob/<release>-historian-prune prune-manual-$(date +%s). Its completion clears the alert. 5. If the PVC has already grown into the warning band, escalate to the Historian Disk-Pressure runbook before DCSNodeRootDiskHigh fires. On a single-node reference cluster, hack/demo-provision.sh recover is the fast path.

DCSNodeRootDiskHigh

Severity: Critical Condition: A node's root filesystem is >75% used for 15min (opt-in rule, default off, enabled via monitoring.prometheusRule.historianDisk.enabled: true).

What happened: Root-fs usage has crossed the NodeHasDiskPressure-precursor threshold. Kubelet's hard-eviction threshold is around 85%. Cleanup pods that need to run during DiskPressure (local-path-provisioner-cleanup, historian-prune) are typically BestEffort and get evicted first once the taint lands.

Actions: 1. SSH to the node and find the offender: ssh root@<node> 'du -h -d 2 /var/lib/rancher/k3s/storage/ /var/log /var/lib/containerd | sort -h | tail -20'. 2. If the historian PVC is the source, check DCSHistorianPruneJobFailing first. That's the upstream cause. Free space without fixing the prune loop and the alert will re-fire within hours. 3. If logs/containerd is the source, run journalctl --vacuum-size=500M and crictl rmi --prune. 4. On a single-node reference cluster, the codified recovery is hack/demo-provision.sh recover (issue #239), idempotent and safe to re-run mid-flight. 5. Production deployments should not see this rule fire at all (it's opt-in and only meaningful on storage without filesystem-quota enforcement). If it does, the deployment is missing a CSI driver (see production-deployment.md § 1).