Skip to content

ADR 0006: Edge-runtime redundancy is hold-then-resume failover to a designated standby node — no hot-standby state replication

Status: Accepted Date: 2026-06-12 Issue: #566

Context

The competitive baseline for a diligence reader is DeltaV, where controller redundancy is a standard orderable product feature: a controller can have a dedicated 1:1 redundant partner receiving continuous state synchronization over a local redundancy link (the documented topologies mount the pair on a shared carrier assembly), giving bumpless switchover. The pairing is per controller, not per unit — DeltaV imposes no controller:unit cardinality, and a controller commonly executes control for many units. Our product today has no controller-level redundancy feature — a failed edge IPC means that unit's control loop is down until the hardware comes back or is replaced (docs/ha-failure-modes.md node-failure row: "manual intervention"). Issue #566 asks Engineering to scope what failover should look like.

What exists today, precisely:

  • One runtime owner per physical unit. The unit runtime is a bare pod (<unit>-runtime) created by the Unit reconciler (internal/controller/physical/unit_pod.go), pinned to a single edge node via UnitSpec.NodeSelector (convention: dcs.io/device=<unit>), hostNetwork: true. There is no lease, leader election, or fencing — the node pin is the one-writer guarantee.
  • Node:unit cardinality is not enforced. Like other control systems, the product is unopinionated about how many units a controller node serves: runtime ports are per-unit configurable (UnitSpec.RuntimeHealthPort/RuntimeGRPCPort), so several runtimes can share a node, though the colliding port defaults make 1:1 the default path. The recommended practice is 1:1 controller:unit — it keeps the failure blast radius to one unit and lets someone unfamiliar with the control system power down one piece of equipment without taking out an unrelated one.
  • Restart-replay autonomy. Every deployed FB program is persisted to hostPath (/var/lib/dcs/runtime/{unit-name}/networks/all.json, NetworkManager.saveAll) and redeployed on pod restart by NetworkManager.RestoreAll, from disk alone; during a control-plane partition the runtime keeps regulating autonomously. A second singleton copy of this state, last-program.json, was read at startup by a path whose writer had been deleted and was removed in #1748. The programs were all that came back until #1776: the commanded values and the block operating points were rebuilt empty, so the first scan after a restart wrote a compile-time default to the field, and the sentence in §Decision.3 below about outputs holding last value across a pod restart was false for as long as it had been published. networks/state.json carries them now and they are installed before the first scan (ADR 0080).
  • Detection and safe-state already exist. A runtime outage is absorbed for a 60 s grace window, after which the phase self-Holds (internal/controller/procedural/phase_controller.go); the unit watchdog issues an ISA-88 Hold when all FIELD drivers are down during an active batch (unit_pod.go; the qualification, and the reason this path could not fire before it, are ADR 0075); Held phases are exempt from the 15 min stuck-phase force-abort. A dead edge node mid-batch therefore lands in Held, not Aborted.
  • Phase logic does not live at the edge. SFC/phase execution runs in the procedural operator (phase_controller.go instantiates the engine), which is already HA via leader election. Mid-chart execution state — active steps, completed steps, variable values, fired transitions — is published to Phase.Status.SFCStatus every second and the engine resumes mid-chart from it (sfc.WithRestoredState). The edge node holds only the FB scan loop and the I/O driver connections.

Three architectural facts shape the option space:

  1. The controller↔I/O binding is a network connection. Our I/O is remote Ethernet I/O (Modbus TCP, OPC UA, EtherNet/IP) — a TCP client connection from the runtime, with no physical attachment between the controller node and the I/O hardware. Moving a unit's runtime to different hardware is therefore a software-only operation: any enrolled node with reach into the unit's field network can take the role, with no rewiring and no physical adjacency requirement on the standby.
  2. Partition autonomy and automatic failover are mutually exclusive per unit. Kubernetes cannot distinguish a dead node from a partitioned one, and our runtime is designed to keep controlling while partitioned. Auto-failover without fencing means two writers on the same field devices — on Modbus, for example, there is no session exclusivity at all; last write wins, undetected.
  3. ISA-88 already defines the recovery shape. Part 1 Clause 7.4 lists control equipment malfunction as a canonical exception event; the standard's response is HOLD (bring equipment to a known safe state) and RESTART with recipe-defined restarting logic. The standard does not ask for invisible failover, and batch processes tolerate Hold/Restart by design.

Decision

Edge-runtime redundancy is hold-then-resume failover: re-binding a unit's runtime to a deployment-designated standby node, with the one-writer guarantee protected by an explicit per-unit availability policy. The product builds the re-binding and fencing mechanism; it does not replicate runtime state to a standby and does not claim bumpless switchover. Which nodes are eligible standbys — and whether a standby is dedicated to one primary or shared — is deployment configuration, not product opinion (product-vs-deployment split).

The scoped mechanism:

  1. Failover targets — UnitSpec.availability names the eligible standby node(s) (explicit list and/or label selector). Targets are enrolled nodes (ADR 0004 adoption contract) with field-network reach to the unit's I/O. The recommended pattern is a dedicated standby per primary controller node: a deterministic failover target is easier to qualify, and it preserves the 1:1 controller:unit practice. A spare set shared across several units is equally supported as a cost option; its consequences are the deployment's to accept consciously — after multiple failures, units can co-locate on one node. The re-bind operation surfaces a co-location guard: it warns (or refuses, per policy) when the target already hosts another unit's runtime.
  2. Unit re-binding as a first-class operation — dcs unit failover <unit> --to-node <node> (and a UI equivalent): the physical operator fences the old binding, rebinds the runtime pod to the target node, and redeploys the control program from the control-plane source of truth (CRDs — no hostPath migration); drivers reconnect to the remote I/O. The action is audit-logged.
  3. Hold-then-resume semantics — the running phase survives in the control plane; the unit is Held during outage and re-bind (the existing watchdog/self-hold path); recovery is an ISA-88 Restart whose recipe-defined restarting logic re-establishes process conditions. Field outputs hold last value (or device fail-safe) during the gap. That was published from the day this was written and was not true of the re-bind itself until ADR 0082: the promoted runtime's first scan wrote a compile-time default, measured on the bench as zero on two live channels four seconds before the unit reported itself controlling. A promoted runtime adopts the operating point the device is holding, and where the device cannot answer it writes nothing at all, so the coupler goes on holding. The comparison to a pod restart is ADR 0080, which is a different mechanism for a different reason: a restart in place remembers, and a promotion must not.
  4. One-writer via availability policy — UnitSpec.availability.mode:
  5. Autonomy (default; today's behavior): on partition the edge keeps controlling. Failover is manual only and requires explicit operator confirmation that the old node is fenced (powered off or disconnected from the field network).
  6. Failover (opt-in): the runtime holds a control lease; on lease expiry it self-fences (stops FB output writes, disconnects drivers), and the operator may auto-rebind to a standby after lease timeout plus margin. Both sides expire against one anchor: every renewal carries the operator's acknowledged-renewal age, and the runtime dates its lease from that age rather than from the arrival of the POST (#579, #1802). A reverse-path (asymmetric) partition, where forward POSTs are delivered and the runtime's responses are dropped, therefore cannot extend the runtime past the operator's own clock, and the same expiry runs the same bounded hold whichever way a partition falls. An epoch/fencing token in the runtime API and write path is defense-in-depth against stale owners.

    The claim above names the wrong fault, and the bench rep says so (#1801, #942 drill 3, 2026-08-24, n=1). The acknowledged-age gate does answer a lost response path, and a reverse-path partition is not one. TCP needs the return path for its ACKs, so a one-way IP drop stalls the forward direction within about a renewal interval: the operator's send window fills and the kernel puts no new data on the wire. On metal exactly one renewal POST landed after the cut, 2.7 s in, and none followed it. The runtime fenced on its own local timer through the ADR 0008 pre-fence hold, which is the symmetric partition's path.

    What the gate does answer is the case where the POSTs keep landing and the operator never gets a usable response back, with the connection still carrying traffic. A wedged or non-answering runtime handler is that case, and TestLeaseBroker_ReportsGrowingAckAgeUnderReversePartition is it in the small: the runtime answers every POST with a 503, the reported age climbs past the lease, and the runtime fences on the report. #579's mechanism is sound. What was published about its coverage was not, and this ADR's own sentence above is where it was published.

    The transport decides something here, and it is not which path fences. With TLS configured the channel is HTTP/2, so the one POST that fits through the closing send window rides a connection that already exists and lands. It carries a fresh acknowledged age, so it extends the runtime's lease past the operator's own clock. Measured, the runtime fenced 10.7 s after the operator had declared the lease Expired, leaving 4.4 s before the re-bind decision where this ADR intends the runtime to have fenced first with the whole margin behind it. Nothing dual-wrote, because the standby took the lease 14.6 s after the fence. With no TLS the channel is HTTP/1.1, the renewal cannot complete a handshake on a cut return path, nothing lands, and the two clocks stay aligned. HTTP/2 is the worse case, which is the reverse of what #1801 assumed before the drill, and the margin it eats is #1802.

    The one-writer decision in this ADR is unchanged, and the extension the drill measured is closed (#1802). The runtime dates its lease from the acknowledged age a report carries rather than from the arrival of the POST that carried it, so both sides run out at one instant and a landed POST buys the runtime nothing. The guarantee rests on one inequality and on nothing else: the re-bind margin is at least the hold bound, and the hold is budgeted from the lease deadline, so a late notice spends the runtime's own safe-state time and never the standby's. No renewal interval and no POST timeout appear in it. "One instant" is one renewal interval loose, and the slack runs in the fail-safe direction (#1810). The age a POST carries is measured before the attempt, so it names the PREVIOUS acknowledged renewal, and the runtime dates its lease from that. The runtime's anchor is therefore always one renewal older than the operator's, and a partitioned runtime self-fences one interval before the operator declares Expired. It has to: the runtime cannot know its own 200 got back, and crediting itself for a response that may be in the bin is exactly what #1802 removed. The inequality above is unaffected, because it is the runtime fencing EARLIER. What the lag does bind is the renewal cadence, which has to be measured against the anchor the runtime holds rather than against the operator's own last success. Measured against the wrong one the two errors cancel, and a single missed renewal retries precisely on the runtime's deadline — at the five-second minimum, half a second past it, which fenced a healthy unit twice off one refused POST. renewInterval takes the anchor's age as its argument for that reason, and TestRenewIntervalAlwaysFitsInsideTheAnchor states the duty with no clock in it.

    TestFailoverMarginCoversHoldBound states it, and TestRuntimeChannelIsHTTP2WhenTLSIsConfigured (pkg/tlsutil) still holds the transport half. That also settled what draining the renewal response body (#1803) does here, and #1803 then measured it rather than assuming it: pooling an HTTP/1.1 connection lets the plaintext posture land its one POST too, and there is no longer an extension for it to inherit. Both postures now fence ahead of the operator's own Expired, and the reverse-path test asserts one outcome across both of them.

The two modes are mutually exclusive by construction — a unit cannot have both partition autonomy and automatic failover. The product ships the mechanism and the safe default; the deployment chooses the policy per unit.

Alternatives Considered

  • Hot-standby state replication with bumpless switchover (DeltaV-style) — a standby runtime continuously receiving FB and variable state from the active one, taking over sub-second. Not chosen: continuous state-sync machinery, a switchover protocol, and a fencing story would be built for a guarantee ISA-88 does not require for batch — Clause 7.4 Hold/Restart is the standard exception path, and batch phases tolerate seconds of held outputs. Note the rejection is of state replication, not of dedicated standby hardware — a dedicated standby per controller is the recommended topology under this ADR; it just stays cold until re-bind. Revisit only if a design partner has a unit where held outputs are process-destructive; that is a new ADR.
  • Plain Kubernetes rescheduling (drop the node pin, let the scheduler move the pod) — not chosen: it breaks the one-writer guarantee on partition. Kubernetes cannot tell dead from partitioned, and the partitioned node's runtime keeps controlling by design, so unfenced rescheduling creates dual writers on live field devices.
  • Status quo (restart-replay only) — not chosen as the terminal posture: detection and Hold already work, but a dead IPC strands the unit until hardware replacement, re-binding is undocumented manual surgery, and the diligence comparison against DeltaV fails.
  • Protocol-level exclusivity as the fencing primitive — not chosen as the universal guarantee, because support varies by protocol. Modbus has no session-ownership concept at all: any client may write, last write wins. OPC UA locking exists only as an optional companion-spec facility (OPC UA DI LockingServices), so it is server-dependent. EtherNet/IP does have a real mechanism — a CIP output assembly accepts a single exclusive-owner connection and rejects a second with an ownership-conflict error (extended status 0x0106) — but ownership is freed after a connection timeout (RPI × timeout multiplier), so under partition it is liveness-bounded rather than absolute, and a recovering stale owner races the new one for re-ownership. The one-writer guarantee therefore lives above the protocol layer (availability policy + lease/epoch fencing); EtherNet/IP exclusive-owner connections are used as defense-in-depth where the device supports them.

Consequences

  • API surfaces that move: UnitSpec gains availability (mode, lease duration, failover targets); the adapter API and FB write path carry a lease/epoch token; the physical operator gains fence, re-bind, and co-location-guard logic; dcs gains a unit failover verb; UI placement to be confirmed before implementation (no inline controls on HMI cards). Alarm + AuditRecord events cover fence, re-bind, and lease loss (21 CFR Part 11).
  • The running phase survives a failover. Phase/SFC state — down to active steps, variable values, and fired transitions — lives in Phase.Status in etcd, not on the edge node; the batch record shows the exception (Hold, Restart) rather than a vanished phase.
  • Hardware guidance changes: reference architectures stop calling the edge node an unmitigated single point of failure and instead document the failover-target patterns — dedicated standby per controller (recommended) vs. shared spares (cost option, co-location trade-off). docs/ha-failure-modes.md, docs/reference-architectures.md (Pattern B), and docs/production-deployment.md update when the mechanism ships; IEC 62443 FR 7 availability targets for the unit runtime can rise accordingly.
  • Marketing posture: until this ships, the accurate claim remains autonomy + restart-replay + supervised Hold. Once shipped, the claim is controller failover (hold-then-resume) to designated standby hardware — never "bumpless redundancy" or redundant-pair parity with DeltaV.
  • Default behavior is unchanged: Autonomy mode preserves today's semantics exactly; everything else is additive and opt-in.
  • Followups: the implementation epic (#574). Warm-standby pre-staging (image pre-pull, a pre-created fenced pod on the designated standby) is a compatible later optimization that shortens failover time without state replication.
  • Reversibility: high before implementation (this is a posture); moderate after — the availability policy field becomes a public CRD contract, and withdrawing automatic failover would be a customer-visible regression requiring a superseding ADR.

Amendment (2026-08-07, #1320): the automatic path issues the Hold itself

Hold-then-resume as implemented waited, before re-binding a Running unit, for "the crash detector or watchdog" to issue the Hold. Both of those live in the pod-lifecycle pass of the same reconcile, downstream of the failover check — so the wait starved the only pass that could end it, and an automatic failover of a unit Running a batch parked forever (found live: eight minutes on a dead node with no Hold, no alarm, no re-bind).

The decision is unchanged; the sequencing within it is corrected. A lease observed Expired means the runtime has self-fenced, which is the crash evidence itself, so the automatic path now issues the ISA-88 Hold directly — through the same code path, alarm, and RuntimeCrashDetected condition as the crash detector — once the expiry outlasts the ADR 0008 hold bound (½ lease, clamped 5–30s). Inside the bound it waits, so a runtime that recovers and renews its lease rides through with no Hold, preserving the crash detector's grace semantics with a lease-sized budget. The re-bind still waits for the safety margin (margin ≥ hold bound by construction), and the lease expiry is still reported (status, alarm, audit record) before any Hold or re-bind. The manual path's refusal of Running units is unchanged.

Amendment (2026-08-31, #1890): start-up is in the lease arithmetic, and the automatic path spends a budget

One control-plane node loss on the bench produced 31 failovers of a controller that never failed, over 23 minutes, and a 1924.6s hole in the process record. It stopped on its own and nothing was done to stop it. ipc-2 was up throughout, and the node that went down carried no runtime and no field path. The same shape had fired three times before, unreported.

A re-bind destroys the thing whose liveness decides whether to re-bind. The newly-bound runtime has to start, pull its image, load its program, connect its drivers and get one renewal acknowledged, and the only window it had for all of that was the lease duration. A lease duration is a renewal budget: it says how long an established runtime may go silent before it has certainly self-fenced, and it is short because it is also how long a dead controller keeps the plant waiting. Spending it on start-up made the mechanism self-sustaining under exactly the conditions that trigger it — a degraded control plane makes the registry, the broker and the field all slow at once. Nothing damped the loop: no backoff, no cap, and no grace for a runtime that had been bound seconds earlier.

It was invisible because every symptom reads as something else. The alarms are failover and lease-expired, which are the correct alarms for a controller that genuinely died. The unit settles Idle on a healthy node, which is a correct-looking resting state. docs/ha-failure-modes.md predicts a failover on lease loss, so the first one matches the documentation exactly. Only the count and the epoch give it away, and nothing surfaced either.

The decision is unchanged. Two things are added to it.

1. The lease has two deadlines, chosen by whether the runtime has ever answered. For a binding this operator has had a renewal acknowledged at, the deadline is unchanged and must stay unchanged: lastSuccess + leaseDuration, which is what the whole one-writer ordering above is built on. For a binding where no renewal has ever been acknowledged, the deadline is availability.runtimeStartupGraceSeconds (default 120s, never shorter than the lease) from the instant the runtime pod first had an address. Expired on that second path is not a detection of a self-fence — no lease was ever granted at that epoch, so nothing is driving anything and there is nobody to fence. It is a verdict that the replacement failed to come up, and it is sized accordingly.

Delaying that verdict can only delay a re-bind, never advance one, so nothing about the one-writer guarantee moves. The asymmetry is what makes a generous default right: too long delays a re-bind for a unit that has no runtime either way, and too short destroys the pod that was seconds from answering.

2. The automatic path may try each free eligible standby once between confirmed runtimes. status.runtimeBinding.unsettledRebinds names the nodes an automatic re-bind has already been sent to since the runtime was last confirmed operating normally. A target in that list has been tried and did not establish, so choosing it again repeats an experiment whose result is in hand. When every free eligible target is named there, the automatic path suspends: it sets FailoverSuspended, raises a Critical alarm of its own — deliberately not under the ordinary failover alarm's prefix, because reading as an ordinary failover is how this hid — and re-binds nothing.

Suspension is a stop on re-binding and not on recovery, and that distinction is what makes it safe. The current binding's pod stays where it is, the broker goes on renewing against it, and the first renewal it acknowledges takes the lease to Held, which clears the condition, the alarm and the ledger through the same edge that clears an ordinary failover's alarm (#1733). In the bench episode that alone would have been the recovery: what prevented it was the re-bind destroying, every 45 seconds, a pod that was 25 seconds into starting. A deliberate dcs unit failover remains available throughout and empties the ledger, because a human choosing a target is a new decision rather than a repetition of this controller's.

A timed automatic retry was considered and refused. It re-enters the same loop at a longer period, and what it retries has already been shown not to work.

3. The expiry says which of the two findings it is. The arithmetic above splits Expired into two events, and for a while only the arithmetic knew. The Info log, the Critical alarm and the Part 11 audit record all went on saying that the runtime had self-fenced and stopped writing. On the never-established path that describes something which did not happen, and it is the sentence an operator reads while deciding whether a controller died. Thirty-one of the bench episode's thirty-two alarms said it about a replacement pod that was still starting.

The broker now captures its verdict at the instant it declares the lease Expired, and all three records read it. An established lease keeps the sentence it had. A lease that was never granted is annunciated as a start-up that failed, naming the grace it was given. The verdict is captured on the transition rather than read afterwards, because the fact that separates the two is whether any renewal has ever been acknowledged, and that is precisely what the runtime finally answering erases.

The verdict also carries what the last renewal attempt got back, and whether the runtime's own handler answered it. A refusal and a lost connection are opposite findings about the same failed renewal (ADR 0081): the refusing runtime is up, reachable and talking, while an unreachable one may be a dead node. Both used to end at V(1), which ADR 0063 makes invisible in a shipped binary. That is why the bench episode's opening expiry could not be explained when this was written: sixty-odd renewals failed against a runtime that was believed to be up, and no durable record anywhere said what a single one of them returned. Part 5 below has the answer, read back off Prometheus.

The audit reason is unchanged on both paths, and deliberately so. The batch-record assembler matches on it to bound the failover data gap (#1308), and a replacement that never came up stops the samples just as surely as a runtime that fenced. What the two paths owe the reader is different prose, not a different correlation key.

4. Whether a runtime has ever answered is a fact about the binding, not about the operator process asking. The two paths above both turn on one question, and the broker answered it out of its own memory. That memory lives in the operator process and dies with it, while the runtime it describes does not: a replacement physical-operator starts a renewal loop against a runtime that may have held the lease for a week and sees exactly what it would see at a pod that has never answered anybody.

Told nothing, the replacement applies both halves of this amendment to an established runtime. It waits out a start-up that finished days ago, which delays the expiry the re-bind margin runs from by the whole of the grace, and then annunciates that expiry as a replacement that failed to come up, about a controller that really did self-fence. Neither is a finding about the plant. Both are the new process's ignorance rendered as one.

status.runtimeBinding.leaseEstablishedBy records it instead. It names the runtime incarnation — pod UID and the runtime container's restart count — that has had a renewal acknowledged at this binding, written once, on the first acknowledgement, by whichever operator observes it. A binding with nothing recorded is one no runtime has ever answered at, which is what a fresh binding looks like and what every re-bind produces, since the re-bind builds the object fresh.

The incarnation and not the pod is what is recorded, for the same reason TerminalStopArming carries the pair. A container that restarts in place keeps its pod, its UID and its binding, and loses every piece of state the lease was established with. That process is starting up and has to earn the lease again, which is precisely the case the start-up budget exists for.

A pass with no acknowledged contact behind it writes nothing. Recording an incarnation this operator has not heard from would hand the next one a claim nobody ever had, and would spend the start-up budget of every fresh binding — which is the storm.

This is the bench episode's own opening loop. A replacement operator took leadership 372 seconds after the node loss and began renewing against ipc-2, bound at epoch 60 throughout, and 30 seconds later declared the lease expired. Why those renewals were not acknowledged is still unanswered, and this changes nothing about that. What it changes is that the answer, when the next occurrence gives one, is no longer read against a runtime the control plane has mistaken for a new one.

5. What the bench episode's two unexplained halves were. Read on 2026-09-03 out of the bench's Prometheus, whose retention reached back past the storm, against T0 = the cord out of cp-1 at 23:50:50Z on 2026-08-31.

The opening expiry was #1918. The runtime on ipc-2 fenced at T0+15 s on its own clock, with the grantor gone. On chart 0.6.3 the liveness probe answered from the same verdict as readiness, so a fenced runtime was NOT_SERVING to both, and the kubelet killed it about a minute into the fence. A Failover runtime starts fenced, so the restarted container was NOT_SERVING from its first probe and was killed again. The restart counter read 1, 2, 3, 4 and 5 at T0+80 s, +140 s, +200 s, +260 s and +320 s, and kube-state-metrics reported the container in CrashLoopBackOff from T0+380 s to the re-bind at T0+422 s. The replacement operator's broker renewed from T0+377 s. Every one of its renewals was a TCP connect refused by a node whose container the kubelet was holding in backoff, which is the transport class of ADR 0081. "ipc-2 was up throughout" was true of the node and false of the process.

The 31 cycles after it were #1933. The broker's persistent volume is pinned to cp-1, so its replacement pod sat Pending from T0+380 s until cp-1 returned at T0+1760 s. A fresh unit-runtime blocks inside Start on the MQTT connect, with no timeout, and opens neither its health server nor its lease endpoint until the broker answers. Each of the 31 pods was created, blocked there, and was deleted by the next re-bind 45 s later. None reached Ready and none restarted. The pod created at T0+1788 s, 28 s after the broker scheduled, is the one that took the lease.

Neither half changes this amendment. The start-up grace is still the right deadline for a runtime that has never answered, and the budget is still what stops the loop. On that night it would have stopped after two re-binds instead of 31, and the pod it left in place would have taken the lease the moment the broker did. What the two halves add is the two causes of a suspension that this ADR's runbooks did not list: a kubelet restarting the runtime the grantor is waiting for, and a broker the replacement cannot reach.

Amendment (2026-09-01, #1894): the grantor keeps granting when it loses the apiserver, and the operator that replaces it waits

Cutting one control-plane node on the bench fenced a unit on a node nobody touched. The node that was cut held the Talos L2 API VIP, the unit's runtime was on an edge node, and the operators were on a third node that stayed up throughout. Nothing the unit depends on was in the blast radius. It self-fenced at T0+29.8s, its lease duration to within the sample interval, drove both outputs to zero and stayed fenced for 37 seconds.

The chain is short and every link is behaving as designed. The API at the VIP was unreachable for 57.615s while the address floated to another node. Every operator pod reaches the apiserver through the in-cluster Service, and every operator whose connection was pinned to the lost endpoint timed out for the whole of that window (the bench verification below is what put the word "pinned" in that sentence). controller-runtime terminates the manager when leader-election renewal fails, which is the correct behaviour for a reconciler: a controller that cannot reach the apiserver cannot be trusted to be the only writer of the objects it owns. The lease broker lives inside physical-operator, so the process exiting took the plant's grantor with it, and the 30s lease expired under a runtime that never moved.

It is the wrong behaviour for a liveness grantor, and the reason is the whole amendment. A grant of liveness is carried on a direct HTTP channel to the runtime's pod IP and never touches the apiserver. Losing the apiserver removes this process's power to move a unit. It does not remove its ability to keep one alive, and it does not remove the duty. The two were tied together only because they happened to share a process.

This is not the failure #1893 answers and the two must not be read as one. There the node carrying the operator is lost, the clock in front of it is the 300s pod-eviction timeout, and a second replica is the answer. Here the apiserver is lost, the clock is a VIP failover measured at 57.6s, and a second replica helps only by chance. The Service address is shared, but each pod's connection is pinned to one apiserver endpoint, so the follower rides the outage through when it is pinned elsewhere and dies with the leader when it is not. Acquiring the leader lease is itself an API write. A graceful drain of the same node removes the API outage entirely and hands every lease off within ±3s, and the unit still expires a lease and still fences, for 6.4s. Only the duration changes.

Three things are added to the decision.

1. A physical-operator whose manager stops without being asked to stands down instead of exiting. For physicalOperator.leaseStandDown (default 90s) both edge liveness brokers go on renewing what they already hold. They learn no new targets, re-bind nothing and write nothing, because the reconcilers that would do any of that stopped with the manager. Then every loop is cancelled and the process exits. A runtime whose control plane came back never notices; a runtime whose grantor never returns fences at the end of the window, which is the outcome the fence exists for. The window is sized to outlast the platform event that takes the apiserver away and then gives it back. That event has a ceiling: a Talos L2 VIP is held through an etcd concurrency session at etcd's default 60s TTL, so an abrupt loss of the holder costs at most what is left of that lease (#1905), and every healthy-quorum figure the bench has produced sits just under it. 90s clears the ceiling by 30s. The first draft of this amendment called it "the measured 57.6s with half of itself again", which was one sample and headroom. A shutdown signal ends it early, because a pod being deleted deliberately is not a pod that lost the API.

The process keeps answering /healthz while it stands down, which is a mechanism detail with a safety consequence. The manager serves that endpoint and has just stopped serving it, so without something answering, the kubelet's liveness probe kills the container most of the way through the very window it was given.

2. An operator that has just taken leadership does not read a lease expiry as a fence. This is the price of the stand-down and it is paid by whichever process takes over. The dangerous shape is a partial partition, where the old leader keeps reaching the runtime while the new leader cannot: the new leader watches its own renewals fail, calls the lease Expired, and re-binds onto a standby while the old runtime is still being held alive. Two writers, which is the one thing this ADR exists to prevent.

The arithmetic that closes it uses only quantities both sides know. The old leader gives up leadership at some instant T, and a challenger cannot acquire before T, because leader election hands the lease over strictly after the holder's renew deadline has passed. The old leader stops renewing at T + leaseStandDown, and the runtime it was holding fences at most one lease duration after that, plus the bounded ADR 0008 pre-fence hold the re-bind margin already covers. So a new leader that waits leaseStandDown + lease duration + margin from its own acquisition has waited past that fence, whatever the partition did to the two of them.

The hold-off gates every re-bind whose fence evidence is the lease expiry, on the automatic road and the manual one alike. It does not gate a re-bind a human has certified with dcs.io/failover-confirm-fenced: powering a node off is a stronger fence than any arithmetic here can construct, and it is asserted about the plant rather than about a clock. The wait is written onto the unit as a FailoverRequest condition with reason LeadershipHoldOff, because a wait nothing renders is the defect #1198 fixed. The condition names the deadline rather than the remainder, and it is the one reason on that condition the operator removes on its own, on the first pass after the deadline with no re-bind in between (#2148). Every other reason is an outcome and stands until the next outcome. This one is a wait, and the bench read "blocked for a further 1m29s" ten days after the wait had ended.

Both halves travel on one value. Setting physicalOperator.leaseStandDown to zero disables the stand-down and the hold-off together, which is the pre-#1894 behaviour: losing the apiserver fences every Failover unit at its lease duration and latches every Autonomy one.

3. A physical-operator that is stopped on purpose hands its leadership straight back. The stand-down answers the abrupt cut and does nothing at all for the graceful one, because a pod being deleted deliberately is not a pod that lost the apiserver, and the two must not be confused. The graceful drain nonetheless fenced a unit for 6.4 seconds with the apiserver never away and every operator lease handed over within ±3s, so something in that path was still too slow, and it was the handover itself.

LeaderElectionReleaseOnCancel was off, which is the controller-runtime scaffold's default. A leader shut down on purpose therefore kept its lease until the full LeaseDuration ran out, so a standby could not begin syncing its caches for fifteen seconds and only then did its broker start renewing unit leases. Against a 30s unit lease that arithmetic has almost nothing left in it, and on the bench it had nothing.

It is now on, for this operator and not for the other four. This is the only one whose absence reaches the field, so it is the only one whose leader-handoff latency is charged against a plant-side budget rather than against reconciliation lag. The scaffold leaves the option off with a warning, and the warning is exactly the thing this ADR now has to keep true: it requires the binary to end immediately when the manager is stopped. Since point 1 this binary does not always end immediately. The two are compatible only because they are disjoint on precisely the condition the warning is about — a release happens when the manager was cancelled, and the stand-down runs only when it was not — so the branch and the option have to be read together and changed together.

Read as one sentence: the authority is handed over as fast as possible when it is being given up on purpose, and held for as long as is safe when this process is only out of touch.

Bench verification (2026-09-02, chart 0.7.1) and two corrections

Three drill 12 reps on the first release carrying both commits, every one an abrupt cord pull at the PDU. Rep 2 is the verification of point 1: the leader, pinned to the cut node's apiserver, lost leadership at T0+14s and stood down; no lease expired and no unit fenced; and the incoming leader was renewing the unit's lease six seconds before the standing-down one released it, so the handover overlapped rather than gapped. Point 3 and the plant cost were measured the same evening, once the bench had a cut phase long enough for the cord to come out with the loop live (every earlier rep that triggered the fix had its field half void, because the demonstration phase's 120s dwell left a 20s cue window). A graceful drain of the node carrying the leader and the VIP moved the grantor lease at T0+1s, expired nothing and fenced nothing, AO held and PV band 59.998 throughout: point 3 verified. An abrupt cut of the node carrying the VIP and the leader's pinned apiserver, leader on a third node and standby pinned elsewhere, engaged the stand-down at T0+15s, had the standby renewing at T0+24s, the API back at T0+59.0s, and the old process releasing at the end of its window at T0+105s with nothing left to regain; no expiry, no fence, AO held at 19628, PV in band. The cost of the fault to the plant was zero on both reps.

First correction: the discriminator is the pinned endpoint and not the VIP. Rep 1 cut the VIP holder and verified nothing, because the leader never lost its API. On the same surviving node the control-operator hit the exact signature at T0+9s and exited while the physical-operator beside it never noticed. In-cluster pods do not use the VIP. The Service address is DNAT'd to a node address, and a pod's long-lived connection stays pinned to that endpoint until the connection dies. So an abrupt control-plane loss reaches roughly one pod in three, whichever node holds the VIP, and the amendment's claim that a second replica "is no answer at all" was too strong. It is an answer two times in three, and the stand-down is the answer the third time. A rep that wants to exercise the stand-down reads the leader's conntrack entry first and cuts the node it names.

Second correction: the window is sized against a ceiling. Rep 3 cut the node holding the VIP and the leader's pinned endpoint together, the stand-down engaged at T0+15s and released at T0+105s as designed, and the unit fenced anyway, for 165s, because the API was away for 307s. That reading is not evidence about the window. cp-1's etcd had been dead since rep 2 with a corrupted store, so cutting cp-2 was the second fault of two and took quorum with it, and without quorum the VIP cannot move at all (#1905). A VIP failover on a healthy quorum is bounded at 60s by the etcd session TTL that elects the holder, and 90s clears it. A quorum loss is bounded by nothing, and a plant that would rather ride one through with no failover capability raises the chart value.

Why the window stays a timer. Rep 3 raised the question of whether the stand-down should end on a condition instead, since the grantor let go with a working channel to the runtime still in hand. It cannot. Point 2's arithmetic is finite only because point 1's window is: the hold-off exists for the partition where the old grantor reaches the runtime and the new leader does not, and in that partition neither "another operator has taken over" nor "I can no longer reach this runtime" ever arrives at the old grantor. With no bound on its tenure there is no instant at which the new leader may read an expiry as a fence, and a Failover unit whose runtime really has died is then never re-bound on the automatic road. The timer is the bound, and the chart value is where a site chooses it.

Amendment (2026-09-02, #1909): the grantor gap is sized against the budget after a renewal

The #1893 second replica was bench-read the same day it shipped, and the arithmetic it was written against did not hold. Drill 12 rep c cut the node carrying the physical-operator leader, with the standby pinned to its own apiserver so nothing but the grantor's death was in the blast radius. Leader election did what was predicted: the standby acquired at T0+16.7s against a 15s election lease polled every 2s, and the plant's spare grantor was back at T0+101s against 397s for the single-replica procedural-operator on the same cut. The unit fenced anyway, for 9.5s, from T0+11.9s to T0+21.0s, with the batch idle so the field cost went unmeasured.

Why the fence landed at +11.9s and not at +30s. The 2026-08-24 amendment above anchors the runtime on the renewal the operator had acknowledged, which is the one BEFORE the POST that carries the report. After a renewal at t the runtime's deadline is therefore t minus the renewal interval plus the lease, and the budget a replacement grantor has to fit inside is the lease minus the interval, less another interval depending on where in the cadence the cut landed. At a third of the lease that was 10 to 20s of a 30s lease. On the rep the last renewal was at T0-8.5s and the fence at T0+11.5s, to the sample. The standby needed 16.7s to acquire and 4.3s more to its first renewal, which is the Unit reconciler reaching reconcileAvailability through everything a Unit reconcile does first. So a two-replica grantor fenced the unit on EVERY abrupt loss of the leader's node, for between one and ten seconds.

The anchor is not the defect. It is what keeps the operator's Expired decision and the runtime's self-fence on one instant, and it is why #1890's re-bind cannot land on a runtime that is still driving. Moving the runtime's anchor to the arrival of the POST would put the fence ten seconds after the operator's own expiry and reopen the two-writer hazard this ADR exists to close. The grantor gap is what has to shrink, and the budget it has to fit inside can be widened without touching the anchor. Three changes, and the decision is that all three ship together, because no one of them alone closes it:

  1. The election is this operator's own. controller-runtime's 15s/10s/2s is a scaffold default for a reconciler, where a lost election is a restart and nothing in the field is timed against it. On the physical-operator the election is the whole of the grantor gap. It runs at 6s/4s/1s (LeaderElectionLeaseDuration and its two siblings, carried by the --leader-election-* flags and physicalOperator.leaderElection in the chart), which puts acquisition at about 8s after an abrupt loss. The 2026-09-01 amendment is what makes the shorter renew deadline affordable: a leader that loses a renewal stands down and goes on granting what it holds, so a spurious loss costs one leader change and nothing in the plant. The other four operators keep the scaffold default.

  2. Every liveness loop starts from the cache on acquisition. A leader-election Runnable (LivenessWarmStart) lists the Units the cache already holds the instant this process wins, and starts the lease and heartbeat loops for every runtime pod the reconciler would have started them for. It builds each target with the same two functions the reconciler calls (leaseTargetFor, heartbeatTargetFor), so the two roads cannot decide differently about one pod, and it writes nothing. The 4.3s becomes one cache list. It is behind the election on purpose: a follower granting liveness would be granting to runtimes it has no authority to move.

  3. The renewal cadence is a sixth of the lease. renewInterval was duration/3, which made the budget after a renewal two thirds of the lease. At duration/6 it is five sixths, 25s of a 30s lease, and 20s in the worst phase. The cost is one POST per unit every 5s instead of every 10s. The #1810 clamp and its tests are unchanged, because the rule they state is about the anchor and not the divisor.

Together: a worst gap of about 10.4s (6s election, two jittered 1s polls, 2s budgeted for the warm start) against a worst budget of 20s. TestGrantorGapFitsInsideTheBudgetAfterARenewal holds the three constants against each other with 5s to spare and fails if any one of them reverts, and deploy/helm/cloud-native-dcs/tests/leader-election.sh holds the chart's rendered values against the same inequality, because a values.yaml edit reaches no Go test.

What this amendment does not claim. The fix is unverified on the bench. The rep that proves it is the same rep c, aimed the same way, and it reads three things the first rep could not: the standby's acquired to its first starting lease renewal loop, which is the warm start's real figure and replaces the 2s budgeted above; the runtime's own control lease expired line minus its leaseDuration, which is the anchor the fence ran from and the only record of the last renewal once the dead leader's log has gone with its node; and whether the unit fenced at all. A fence of any length on that rep reopens this amendment.

A site whose apiserver cannot answer a Lease update inside four seconds widens all three election timings together, and pays for it in the budget above. client-go refuses a set where the lease does not exceed the renew deadline, or the renew deadline does not exceed the retry period with its jitter, and it refuses at manager construction, so a misconfigured operator exits before it can grant anything.

Amendment (2026-09-02, #1918): a fence is not a death, so the two probes ask two questions

Drill 12 rep 3 of the #1894 verification fenced a unit for 165s. That was a quorum loss (#1905) and the fence itself was correct. Inside it the runtime container exited and was restarted by the kubelet:

19:01:38.253  control lease expired; running bounded edge-local hold before self-fencing
19:01:38.362  control lease expired, self-fenced: FB output writes stopped, reads continue
19:02:20.417  shutting down unit runtime
              (exit 1; container restarted 19:03:51)

The pod's liveness and readiness probes both asked the same gRPC health service on 61052, and the runtime answered both with one verdict. A fenced runtime answered NOT_SERVING, which is right for readiness and was fatal for liveness: three failures at a 20s period landed the restart about a minute into the fence. Every fence shorter than that was unaffected, which is why #1909's 9.5s and the original 37s never showed it.

A fence is the runtime doing the one thing this ADR asks of it. It is not an unhealthy process. Restarting it inside the fence throws away the process that ran the hold program, forces the self-held snapshot through the #1804 recovery road, and on the bench opened the historian gap at the same second (Data gap opens vs T0=+165.950s against a restart at T0+166s). The 165s fence in that rep was two events, a fence and a restart, and the row recorded one. The #1894 stand-down exists so a runtime survives a control-plane absence untouched, and this was the kubelet touching it on a schedule.

The same verdict had a second victim that the bench did not have to show. Every field driver gone past the 60s grace was also NOT_SERVING, so the liveness probe restarted a runtime for losing its plant. A restart reconnects no cable. The reconnecting driver is the in-process remedy, and the restart threw away the held state while it ran.

The decision. Readiness and liveness are different questions, and the gRPC health protocol can answer several named services from one server, so they are two services (pkg/grpc: ServiceReadiness, ServiceLiveness).

  • Readiness (readiness, and the protocol's unnamed overall status) is what it has always been: NOT_SERVING while fenced, and NOT_SERVING when every field driver has been gone past the grace period. The unit controller reads pod readiness to see the outage, and that path is unchanged.
  • Liveness (liveness) answers whether the process is alive and scanning. Fenced, held and degraded are all SERVING. The one NOT_SERVING it can give is a network that reports itself Running or Degraded and has not completed a scan in longer than its stall budget, 60s or ten scan intervals, whichever is longer. A wedged scan goroutine holds the plant's last commanded state and regulates nothing, nothing inside the process can free it, and a restart is the remedy. That is the only case where it is.

The health server refuses the liveness service with NotFound when it was built without a liveness verdict, and refuses any name it does not know. A probe naming the wrong service therefore fails loudly on the first probe, never quietly on the readiness verdict under another name.

What holds it. The unit runtime's pod is built by the physical-operator and not by the chart, so the gate is TestUnitPodProbesNameDifferentHealthServices in internal/controller/physical: the liveness probe names ServiceLiveness, the readiness probe names ServiceReadiness, and the two constants differ. It was proved by mutation: pointing both probes at the readiness service fails it on the liveness line. TestLivenessCheck_FencedRuntimeIsNotReadyAndAlive holds the two verdicts apart on a fence, and TestLivenessCheck_StalledScanIsDeath holds the one case where liveness does fail.

What this amendment does not claim. The restart was read out of one container log, and the fix is unverified on the bench. The rep that proves it is any fence longer than about 90s with the runtime's restart count read before and after. The row in ha-failure-modes.md that said the kubelet restarts the pod after all drivers are lost was bench-validated on the Hold at +21.1s and not on the restart, and it no longer says so.

Amendment (2026-09-03, #1930): a dead apiserver connection is found by the ping, and the ping is timed against the election

The #1909 amendment was bench-read the day after it shipped, on chart 0.7.3, drill 12 rep c: cp-3 cut at the PDU with the physical-operator leader and the VIP on it, the standby on cp-2. The three fixes did what they shipped to do. From the moment the standby could reach an apiserver it acquired within the 6s election lease, and acquisition to starting lease renewal loop was 1.4s against the 2s budgeted. The unit fenced anyway, for 29s, from T0+23s to T0+52s, because for 43s the standby had no apiserver at all. Twelve lease reads in a row timed out at 2s each, and then at T0+43s an unrelated audit write logged http2: client connection lost, the standby dialled again, and the election ran.

The 43s was client-go's, not the node controller's. The issue as filed read the standby's recovery as the kubernetes Endpoints being pruned when cp-3 went NodeNotReady at T0+44s. That was a coincidence of two clocks. The apiservers prune each other by lease, 15s TTL reconciled every 10s, and node status has no hand in it. What held the standby was its own connection. Every in-cluster client dials the Service address, kube-proxy DNATs the connection to one apiserver endpoint, and client-go keeps that HTTP/2 connection for as long as it looks alive. A node losing power sends no FIN and no RST, so the connection looked alive, and every request rode it until the transport's health check gave up: a PING after 30s with no frame received, a close after 15s more with no answer. The standby's last frame from cp-3 arrived a moment before the cut, and 43s later the connection was declared lost. #1894 had found the same pin deciding the LEADER's fate and answered it with the stand-down. The standby has nothing to stand down from. It can only wait for its transport, and its transport was waiting on a 45s clock nobody had put in the arithmetic.

The decision. The two timeouts are client-go's own, read from the process environment (HTTP2_READ_IDLE_TIMEOUT_SECONDS, HTTP2_PING_TIMEOUT_SECONDS) when a transport is built and from nowhere else, so the physical-operator sets them before it builds one: 2s and 2s (APIServerReadIdleTimeout, APIServerPingTimeout, flags --apiserver-read-idle-timeout and --apiserver-ping-timeout, chart physicalOperator.apiserverHealthCheck). Detection inside four seconds of the last frame, in place of forty-five.

The detection does not overlap the election lease, and the arithmetic has to say so. A challenger restarts its lease clock the first time it reads a record it has not seen, because it cannot know how long ago the holder wrote it. The leader renews every second and the standby reads every second, so a renewal between the standby's last read and the cut is the usual case, and the first successful read after the reconnect is that unseen record. The worst gap is therefore the detection, then a poll, then the whole election lease, then a poll, then the warm start: 4 + 6 + 2.4 + 2, about 14.4s, against the 20s the runtime has after a renewal in the worst phase, with 5s to spare. TestGrantorGapFitsInsideTheBudgetAfterARenewal carries the fourth term now and fails at client-go's defaults, where the sum is 55s. deploy/helm/cloud-native-dcs/tests/leader-election.sh holds the chart's rendered values against the same sum, refuses a read-idle of zero (client-go's documented disable, which restores the 43s) and a fraction of a second (which client-go would read as something else), and clears the block to prove the flags are the values and not the template.

TestAPIServerHealthCheck_ADeadConnectionIsGivenUpInsideTheWindow is the mechanism itself, driven through rest.HTTPClientFor against a relay that goes silent without closing, which is what a node losing power looks like to its peer. The tuned client re-dialled 2.02s after the peer went silent at 1s and 1s. The control case, the same fault under client-go's defaults, kept the dead connection for the whole window, which is the bench's 43s reproduced and is what proves the environment moved it rather than the relay.

The price. One PING per connection after two quiet seconds, and a fresh TCP and TLS handshake when an apiserver cannot answer a PING inside two more, with the process's watches resuming on the new connection. On the leader a spurious close can cost a renewal, which since the 2026-09-01 amendment costs one leader change and nothing in the plant. What it can never cost is a grant of liveness, which travels to a runtime pod IP and never touches the apiserver. The other four operators keep client-go's defaults: a reconciler that sits on a dead connection for 45s restarts on leader election lost, and nothing in the field is timed against it.

What this amendment does not cover. A fresh dial while the dead endpoint is still in the kubernetes Endpoints. For up to 25s after the cut, until the apiservers' lease reconciler drops it, a new dial lands on the dead node one time in three on a three-node control plane and hangs for the election client's 2s request timeout before the next poll dials again. Each such dial is one term of the 5s margin, so two in a row spend it in the worst phase and three fence the unit, which is one rep in twenty-seven at the worst phase and fewer elsewhere. The issue's alternative closes that residue too, at a cost this amendment does not take: dialling the node-local apiserver over status.hostIP gives the standby a connection that cannot die with the leader's node, and it requires every physical-operator pod to be scheduled on a control-plane node, which the reference bench does and a production install need not. That stays an opt-in for the founder to rule on, not a default.

What this amendment does not claim. The fix is unverified on the bench. The rep that proves it is rep c with the standby's pin READ as the leader's node before the cut, which the census could not do at the cue because a standby's 2s election reads slip between conntrack samples. The reading is the standby's own log: its last Error retrieving lease lock to Successfully acquired lease should be under 10s, the runtime should log no control lease expired, and dcs_runtime_fenced should stay at 0. A fence of any length on that rep reopens this amendment. Rep c as the #1909 amendment specifies it, standby pinned to a surviving apiserver, is a different rep and is still owed.

Amendment (2026-09-03, #1933): a broker is a transport, and a runtime starts without one

The 2026-08-31 storm had two halves the #1890 amendment did not explain, and both were read back off the bench's Prometheus on 2026-09-03. kube-state-metrics was scraped throughout, with a week of retention. The opening expiry was #1918: five liveness kills of a fenced runtime, then CrashLoopBackOff across the replacement grantor's renewal window. The 31 cycles that followed were this defect.

What the runtime did. Adapter.Start connected the device drivers non-fatally and then called the MQTT client's blocking connect with the process context, which is cancelled on SIGTERM and by nothing else. autopaho's AwaitConnection returns when a connection is up or the context is done. cmd/unit-runtime opens the gRPC health server and the HTTP API only after Start returns. So for as long as its broker was unreachable a fresh runtime listened on nothing: no readiness, no liveness, and no lease endpoint. A Failover runtime could not be granted its control lease, and it could not say so. An established runtime was never affected. Its client reconnects in the background, the store-and-forward queue holds its publishes, and its lease keeps renewing. The defect was on start-up only, which is exactly the phase the #1890 arithmetic measures.

What it did on the bench. The broker's PersistentVolumeClaim is local-path, so it pinned the broker to cp-1, the node that was cut. The replacement broker pod was created at T0+380s and stayed Pending until cp-1 came back at T0+1760s. Between T0+422s and T0+1788s the automatic path created 31 runtime pods, one every 45.5s, alternating between the two device nodes. None reached Ready and none restarted. Each was a container blocked in Start, deleted by the next re-bind. The pod created at T0+1788s, 28s after the broker was scheduled, took the lease at T0+1805s. The storm ended when the broker did, and nothing else changed at that minute. The healthy re-bind on 2026-08-27 had taken 10s to Held with the broker up.

The decision. A broker is a telemetry and audit transport. It is not a precondition of control, and the start path now says so the way the hold path has since #1779. Start opens the connection with the client's ConnectBackground and returns. The error that call can return is a configuration error, an unparseable URL or unreadable TLS material, and never reachability. The retained status a starting runtime publishes rides the store-and-forward queue while the broker is away, in order ahead of anything published later, and the replay worker's on-connect listener delivers it when the broker answers. The health server and the lease endpoint therefore come up on a node whose broker is away. Readiness still answers to the lease, as it did.

A returning broker is also told what the runtime is now. The queue can evict the start-time status in a long outage, and the broker that comes back can be a fresh one holding no retained state. So every connect re-announces the current status, and the announcement goes through the queue rather than around it. A live publish from the connect listener would race the replay of an older status still queued, and the older one landing second would leave the broker asserting a fence state the runtime had already left. The queue is FIFO, so the announcement is delivered after everything that was true before it.

The mqttPublisher slice the runtime holds carries ConnectBackground and not Connect. The blocking form is the one call in the client that can hold a goroutine for the length of an outage, and the runtime has no goroutine it can spend that way. Leaving it off the interface is what keeps it from coming back.

The tests. TestStart_ALeaseIsGrantedWhileTheBrokerIsUnreachable starts a Failover runtime against a closed loopback port, bounds Start at ten seconds, grants the lease over the HTTP endpoint with the client still reporting disconnected, and finds both status publishes in the queue. With the blocking connect restored it fails at the bound, so the bound is the assertion. TestStart_TheStatusIsAnnouncedOnEveryConnect brings the fake transport up, grants the lease, takes the transport down and brings it up again, and reads the last status on the topic each time. Without the connect listener it fails on the first connect.

What this amendment does not cover. The device drivers are still connected inside Start, non-fatally, and a driver whose dial has a long timeout still delays the servers by that timeout. Nothing on the bench has measured that phase. (The same bench reading measured it that afternoon, and the #1936 amendment below answers it.) A runtime with the queue disabled (--mqtt-queue-size 0) publishes its announcement live on a goroutine of its own, because there is nothing for it to race, and it drops the start-time status the way it drops everything else the broker did not take. The bench rep that found this, drills/rebind-budget-rep.sh in cndcs-deploy-bench, reads the replacement Held inside the start-up grace with the broker still down once a chart carrying this amendment is on the rig, and its re-bind budget claims do not apply on that chart. The bench is pinned to 0.7.4, which predates this, so that reading is owed.

Amendment (2026-09-03, #1936): a driver is dialled beside the start, and a dial that failed is still watched

The #1933 amendment closed with the phase it did not cover: the device drivers were still dialled inside Start, one after another, and nothing had measured it. The bench reading that closed #1935 the same afternoon had the numbers in the replacement pod's own log. starting runtime at 17:16:33.215, failed to connect IO driver, will retry on first I/O at 17:16:36.281 with dial tcp 10.10.20.51:502: connect: no route to host, and runtime started successfully three milliseconds later. One Modbus IOModule, on a node with no route to the coupler, cost 3.07s of the 15s to grant. That is the kernel's ARP clock, three solicitations at a second each. A coupler that answers ARP and drops the connection costs the transport's own timeout instead, 5s for Modbus and 10s for OPC UA, and the loop paid it per IOModule in sequence. cmd/unit-runtime opens the gRPC health server and the lease endpoint after Start returns, so a unit with several modules on an unreachable segment would have spent tens of seconds unable to be probed or granted. The same class as #1933 and #1935: a network round trip on the goroutine the control plane depends on, bounded by nothing the caller set.

The decision. A driver on a network protocol is dialled on a goroutine of its own, and Start returns with the dial in flight. The outcome of the dial was never a precondition of anything that follows it. A driver whose dial fails already answers "will retry on first I/O", the scan's first read is that I/O, and the outcome is logged from the goroutine in the same words it was logged in before. The discriminator is the set wrapIfNetwork reads, because it is the ReconnectingDriver the deferral leans on. The simulation driver is connected before Start returns, as it was: its reads refuse until it is connected, nothing re-dials it, and its connect touches no wire.

Two things in the wrapper had to change for the deferral to be safe, and the issue named the first of them as the thing to check before deciding.

The health monitor now starts whether or not the first dial succeeded. It used to start on success only. So a device that was down when the runtime started was dialled again by the next I/O and by nothing else, while a device that dropped after a successful start was re-dialled every HealthInterval. Those are one fault seen from two moments, and with the dial no longer waited on, the first moment is the ordinary state of a replacement runtime on a node that cannot yet reach its coupler. Both are watched the same way now.

The first dial runs under the wrapper's mutex, the way every reconnect always has. A read that arrives while the first dial is on the wire waits for its answer and finds the driver connected, instead of dialling beside it. And a driver the adapter retires, on reload or on Stop, is closed: its Disconnect waits for a dial in flight, and a dial that has not started yet, from Connect or from a late reader's lazy reconnect, is refused with ErrDriverClosed. Without that, a retirement landing before the deferred dial started would have left a driver reconnecting a device nothing reads, with a health monitor behind it that nothing would ever stop.

What the reload path does. reloadIOConfigs still dials a new driver synchronously. It runs on the file watcher's goroutine, which nothing in the control plane waits on, and a test reads the driver connected the instant the reload returns. The issue's subject was Start, and the deferral stays there.

The tests. TestStart_ALeaseIsGrantedWhileTheCouplerHasNotAnswered gives the runtime two IO drivers whose dials never answer and one simulation driver that takes 300ms to connect, bounds Start at ten seconds, reads both dials in flight and the simulation driver connected the moment Start returns, grants the lease over the HTTP endpoint with both dials still on the wire, then lets one coupler answer and reads that driver connected. Restoring the synchronous dial fails it at the bound; deferring the simulation driver too fails it on the 300ms. In the driver package, TestReconnectingDriver_AFailedFirstDialIsStillWatched refuses the first dial and reads the health monitor bring the driver up with nothing else dialling. TestReconnectingDriver_AReadDuringTheFirstDialWaitsForIt holds a dial on the wire, issues a read, and counts one dial and not two. TestReconnectingDriver_ARetiredDriverDoesNotDial retires a driver before its first dial and during it. Eight mutations, one per claim, each reddened the test that carries it.

What is owed. On the next chart carrying this, the bench reading is the gap between starting runtime and runtime started successfully with the coupler unreachable, which should read in milliseconds, with the failed to connect IO driver line arriving after gRPC health server listening.

Amendment (2026-09-22, #2151): a carried lease state names its epoch

The #1890 amendment taught the operator that an Expired is two findings under one state name, and that the one where no renewal was ever acknowledged at the binding is a replacement that failed to come up. Its test started from an empty lease state, and the one road where a replacement exists never reached that verdict. performFailover carries the predecessor's Expired onto the binding it builds, so that the standby's first renewal is an Expired to Held edge and writes the LeaseReestablished record #1308's assembler closes the data gap on. noteLeaseState opened with state == rb.LeaseState. The broker's own Expired for a standby that never answered was equal to the carried one, so the operator wrote no log line, no alarm and no audit record for it. The plant was told the primary had fenced and that a failover was pending, and nothing about the runtime that was tried and never took the lease.

The binding carries leaseStateEpoch beside leaseState now: the epoch the broker observed the state at. A re-bind carries both values unchanged under the new epoch, so a carried state is one whose epoch is behind the binding's, and LeaseStateCarried in the API package is the one reading of that, shared with the gateway. noteLeaseState treats a state equal to a carried one as a transition, records it with #1890's start-up verdict, and stamps the epoch, which is what keeps the next pass from recording it again. A state written before the field existed has no epoch and is read as the binding's own, so an upgrade does not re-record every unit that happened to be fenced at the time.

The audit reason stays LeaseExpired, as #1890 chose for the first-binding case, and the batch-record assembler is taught the consequence: a gap opens once. The second failover's opening expiry is the latest LeaseExpired at or before its anchor walked back to the first of its run, where a run is broken by a re-establishment or a runtime-restored record and bounded by the batch's start. Without the walk the replacement's own record, which carries no acknowledged contact because the replacement never answered, would have moved the second event's opening off the instant the field went quiet and downgraded it to Overlapping with a warning about a lease duration the samples stopped long before.

The alarm is not doubled. The lease-expired alarm dedupes by prefix against the active one, which still describes the primary's self-fence, and that sentence is still true. The replacement's verdict reaches the log, the Part 11 record and the binding, and the gateway now reports Expired for that binding because the broker declared it there, which #2150's reading of leaseEstablishedBy alone could not tell from the carried case.

The tests. TestReconcileAvailability_AReplacementThatNeverComesUpIsRecorded starts from the binding performFailover leaves and a runtime that refuses every renewal, reads nothing recorded before the broker has spoken, one alarm and one LeaseExpired record with the start-up verdict after it has, and nothing more on a third pass. TestReconcileAvailability_ARecordedReplacementExpiryStillReestablishes lets the standby answer late and reads the LeaseReestablished record behind it. TestNoteLeaseState_AStateWithNoEpochIsThisBindingsOwn is the upgrade guard. TestDataGapOpensOnceAcrossAReplacementThatNeverCameUp and TestDataGapReopensAfterAReestablishment hold the assembler in both directions. Restoring the equality check reddens the first and fourth; removing the walk reddens the assembler test.