ADR 0006: Edge-runtime redundancy is hold-then-resume failover to a designated standby node — no hot-standby state replication¶
Status: Accepted Date: 2026-06-12 Issue: #566
Context¶
The competitive baseline for a diligence reader is DeltaV, where
controller redundancy is a standard orderable product feature: a
controller can have a dedicated 1:1 redundant partner receiving
continuous state synchronization over a local redundancy link (the
documented topologies mount the pair on a shared carrier assembly),
giving bumpless switchover. The pairing is per controller, not per
unit — DeltaV imposes no controller:unit cardinality, and a controller
commonly executes control for many units. Our product today has no
controller-level redundancy feature — a
failed edge IPC means that unit's control loop is down until the
hardware comes back or is replaced (docs/ha-failure-modes.md
node-failure row: "manual intervention"). Issue #566 asks Engineering
to scope what failover should look like.
What exists today, precisely:
- One runtime owner per physical unit. The unit runtime is a bare
pod (
<unit>-runtime) created by the Unit reconciler (internal/controller/physical/unit_pod.go), pinned to a single edge node viaUnitSpec.NodeSelector(convention:dcs.io/device=<unit>),hostNetwork: true. There is no lease, leader election, or fencing — the node pin is the one-writer guarantee. - Node:unit cardinality is not enforced. Like other control
systems, the product is unopinionated about how many units a
controller node serves: runtime ports are per-unit configurable
(
UnitSpec.RuntimeHealthPort/RuntimeGRPCPort), so several runtimes can share a node, though the colliding port defaults make 1:1 the default path. The recommended practice is 1:1 controller:unit — it keeps the failure blast radius to one unit and lets someone unfamiliar with the control system power down one piece of equipment without taking out an unrelated one. - Restart-replay autonomy. Every deployed FB program is persisted
to hostPath (
/var/lib/dcs/runtime/{unit-name}/networks/all.json,NetworkManager.saveAll) and redeployed on pod restart byNetworkManager.RestoreAll, from disk alone; during a control-plane partition the runtime keeps regulating autonomously. A second singleton copy of this state,last-program.json, was read at startup by a path whose writer had been deleted and was removed in #1748. The programs were all that came back until #1776: the commanded values and the block operating points were rebuilt empty, so the first scan after a restart wrote a compile-time default to the field, and the sentence in §Decision.3 below about outputs holding last value across a pod restart was false for as long as it had been published.networks/state.jsoncarries them now and they are installed before the first scan (ADR 0080). - Detection and safe-state already exist. A runtime outage is
absorbed for a 60 s grace window, after which the phase self-Holds
(
internal/controller/procedural/phase_controller.go); the unit watchdog issues an ISA-88 Hold when all FIELD drivers are down during an active batch (unit_pod.go; the qualification, and the reason this path could not fire before it, are ADR 0075); Held phases are exempt from the 15 min stuck-phase force-abort. A dead edge node mid-batch therefore lands in Held, not Aborted. - Phase logic does not live at the edge. SFC/phase execution runs in
the procedural operator (
phase_controller.goinstantiates the engine), which is already HA via leader election. Mid-chart execution state — active steps, completed steps, variable values, fired transitions — is published toPhase.Status.SFCStatusevery second and the engine resumes mid-chart from it (sfc.WithRestoredState). The edge node holds only the FB scan loop and the I/O driver connections.
Three architectural facts shape the option space:
- The controller↔I/O binding is a network connection. Our I/O is remote Ethernet I/O (Modbus TCP, OPC UA, EtherNet/IP) — a TCP client connection from the runtime, with no physical attachment between the controller node and the I/O hardware. Moving a unit's runtime to different hardware is therefore a software-only operation: any enrolled node with reach into the unit's field network can take the role, with no rewiring and no physical adjacency requirement on the standby.
- Partition autonomy and automatic failover are mutually exclusive per unit. Kubernetes cannot distinguish a dead node from a partitioned one, and our runtime is designed to keep controlling while partitioned. Auto-failover without fencing means two writers on the same field devices — on Modbus, for example, there is no session exclusivity at all; last write wins, undetected.
- ISA-88 already defines the recovery shape. Part 1 Clause 7.4 lists control equipment malfunction as a canonical exception event; the standard's response is HOLD (bring equipment to a known safe state) and RESTART with recipe-defined restarting logic. The standard does not ask for invisible failover, and batch processes tolerate Hold/Restart by design.
Decision¶
Edge-runtime redundancy is hold-then-resume failover: re-binding a unit's runtime to a deployment-designated standby node, with the one-writer guarantee protected by an explicit per-unit availability policy. The product builds the re-binding and fencing mechanism; it does not replicate runtime state to a standby and does not claim bumpless switchover. Which nodes are eligible standbys — and whether a standby is dedicated to one primary or shared — is deployment configuration, not product opinion (product-vs-deployment split).
The scoped mechanism:
- Failover targets —
UnitSpec.availabilitynames the eligible standby node(s) (explicit list and/or label selector). Targets are enrolled nodes (ADR 0004 adoption contract) with field-network reach to the unit's I/O. The recommended pattern is a dedicated standby per primary controller node: a deterministic failover target is easier to qualify, and it preserves the 1:1 controller:unit practice. A spare set shared across several units is equally supported as a cost option; its consequences are the deployment's to accept consciously — after multiple failures, units can co-locate on one node. The re-bind operation surfaces a co-location guard: it warns (or refuses, per policy) when the target already hosts another unit's runtime. - Unit re-binding as a first-class operation —
dcs unit failover <unit> --to-node <node>(and a UI equivalent): the physical operator fences the old binding, rebinds the runtime pod to the target node, and redeploys the control program from the control-plane source of truth (CRDs — no hostPath migration); drivers reconnect to the remote I/O. The action is audit-logged. - Hold-then-resume semantics — the running phase survives in the control plane; the unit is Held during outage and re-bind (the existing watchdog/self-hold path); recovery is an ISA-88 Restart whose recipe-defined restarting logic re-establishes process conditions. Field outputs hold last value (or device fail-safe) during the gap. That was published from the day this was written and was not true of the re-bind itself until ADR 0082: the promoted runtime's first scan wrote a compile-time default, measured on the bench as zero on two live channels four seconds before the unit reported itself controlling. A promoted runtime adopts the operating point the device is holding, and where the device cannot answer it writes nothing at all, so the coupler goes on holding. The comparison to a pod restart is ADR 0080, which is a different mechanism for a different reason: a restart in place remembers, and a promotion must not.
- One-writer via availability policy —
UnitSpec.availability.mode: Autonomy(default; today's behavior): on partition the edge keeps controlling. Failover is manual only and requires explicit operator confirmation that the old node is fenced (powered off or disconnected from the field network).-
Failover(opt-in): the runtime holds a control lease; on lease expiry it self-fences (stops FB output writes, disconnects drivers), and the operator may auto-rebind to a standby after lease timeout plus margin. Both sides expire against one anchor: every renewal carries the operator's acknowledged-renewal age, and the runtime dates its lease from that age rather than from the arrival of the POST (#579, #1802). A reverse-path (asymmetric) partition, where forward POSTs are delivered and the runtime's responses are dropped, therefore cannot extend the runtime past the operator's own clock, and the same expiry runs the same bounded hold whichever way a partition falls. An epoch/fencing token in the runtime API and write path is defense-in-depth against stale owners.The claim above names the wrong fault, and the bench rep says so (#1801, #942 drill 3, 2026-08-24, n=1). The acknowledged-age gate does answer a lost response path, and a reverse-path partition is not one. TCP needs the return path for its ACKs, so a one-way IP drop stalls the forward direction within about a renewal interval: the operator's send window fills and the kernel puts no new data on the wire. On metal exactly one renewal POST landed after the cut, 2.7 s in, and none followed it. The runtime fenced on its own local timer through the ADR 0008 pre-fence hold, which is the symmetric partition's path.
What the gate does answer is the case where the POSTs keep landing and the operator never gets a usable response back, with the connection still carrying traffic. A wedged or non-answering runtime handler is that case, and
TestLeaseBroker_ReportsGrowingAckAgeUnderReversePartitionis it in the small: the runtime answers every POST with a 503, the reported age climbs past the lease, and the runtime fences on the report. #579's mechanism is sound. What was published about its coverage was not, and this ADR's own sentence above is where it was published.The transport decides something here, and it is not which path fences. With TLS configured the channel is HTTP/2, so the one POST that fits through the closing send window rides a connection that already exists and lands. It carries a fresh acknowledged age, so it extends the runtime's lease past the operator's own clock. Measured, the runtime fenced 10.7 s after the operator had declared the lease Expired, leaving 4.4 s before the re-bind decision where this ADR intends the runtime to have fenced first with the whole margin behind it. Nothing dual-wrote, because the standby took the lease 14.6 s after the fence. With no TLS the channel is HTTP/1.1, the renewal cannot complete a handshake on a cut return path, nothing lands, and the two clocks stay aligned. HTTP/2 is the worse case, which is the reverse of what #1801 assumed before the drill, and the margin it eats is #1802.
The one-writer decision in this ADR is unchanged, and the extension the drill measured is closed (#1802). The runtime dates its lease from the acknowledged age a report carries rather than from the arrival of the POST that carried it, so both sides run out at one instant and a landed POST buys the runtime nothing. The guarantee rests on one inequality and on nothing else: the re-bind margin is at least the hold bound, and the hold is budgeted from the lease deadline, so a late notice spends the runtime's own safe-state time and never the standby's. No renewal interval and no POST timeout appear in it. "One instant" is one renewal interval loose, and the slack runs in the fail-safe direction (#1810). The age a POST carries is measured before the attempt, so it names the PREVIOUS acknowledged renewal, and the runtime dates its lease from that. The runtime's anchor is therefore always one renewal older than the operator's, and a partitioned runtime self-fences one interval before the operator declares
Expired. It has to: the runtime cannot know its own 200 got back, and crediting itself for a response that may be in the bin is exactly what #1802 removed. The inequality above is unaffected, because it is the runtime fencing EARLIER. What the lag does bind is the renewal cadence, which has to be measured against the anchor the runtime holds rather than against the operator's own last success. Measured against the wrong one the two errors cancel, and a single missed renewal retries precisely on the runtime's deadline — at the five-second minimum, half a second past it, which fenced a healthy unit twice off one refused POST.renewIntervaltakes the anchor's age as its argument for that reason, andTestRenewIntervalAlwaysFitsInsideTheAnchorstates the duty with no clock in it.TestFailoverMarginCoversHoldBoundstates it, andTestRuntimeChannelIsHTTP2WhenTLSIsConfigured(pkg/tlsutil) still holds the transport half. That also settled what draining the renewal response body (#1803) does here, and #1803 then measured it rather than assuming it: pooling an HTTP/1.1 connection lets the plaintext posture land its one POST too, and there is no longer an extension for it to inherit. Both postures now fence ahead of the operator's own Expired, and the reverse-path test asserts one outcome across both of them.
The two modes are mutually exclusive by construction — a unit cannot have both partition autonomy and automatic failover. The product ships the mechanism and the safe default; the deployment chooses the policy per unit.
Alternatives Considered¶
- Hot-standby state replication with bumpless switchover (DeltaV-style) — a standby runtime continuously receiving FB and variable state from the active one, taking over sub-second. Not chosen: continuous state-sync machinery, a switchover protocol, and a fencing story would be built for a guarantee ISA-88 does not require for batch — Clause 7.4 Hold/Restart is the standard exception path, and batch phases tolerate seconds of held outputs. Note the rejection is of state replication, not of dedicated standby hardware — a dedicated standby per controller is the recommended topology under this ADR; it just stays cold until re-bind. Revisit only if a design partner has a unit where held outputs are process-destructive; that is a new ADR.
- Plain Kubernetes rescheduling (drop the node pin, let the scheduler move the pod) — not chosen: it breaks the one-writer guarantee on partition. Kubernetes cannot tell dead from partitioned, and the partitioned node's runtime keeps controlling by design, so unfenced rescheduling creates dual writers on live field devices.
- Status quo (restart-replay only) — not chosen as the terminal posture: detection and Hold already work, but a dead IPC strands the unit until hardware replacement, re-binding is undocumented manual surgery, and the diligence comparison against DeltaV fails.
- Protocol-level exclusivity as the fencing primitive — not chosen
as the universal guarantee, because support varies by protocol.
Modbus has no session-ownership concept at all: any client may write,
last write wins. OPC UA locking exists only as an optional
companion-spec facility (OPC UA DI
LockingServices), so it is server-dependent. EtherNet/IP does have a real mechanism — a CIP output assembly accepts a single exclusive-owner connection and rejects a second with an ownership-conflict error (extended status 0x0106) — but ownership is freed after a connection timeout (RPI × timeout multiplier), so under partition it is liveness-bounded rather than absolute, and a recovering stale owner races the new one for re-ownership. The one-writer guarantee therefore lives above the protocol layer (availability policy + lease/epoch fencing); EtherNet/IP exclusive-owner connections are used as defense-in-depth where the device supports them.
Consequences¶
- API surfaces that move:
UnitSpecgainsavailability(mode, lease duration, failover targets); the adapter API and FB write path carry a lease/epoch token; the physical operator gains fence, re-bind, and co-location-guard logic;dcsgains aunit failoververb; UI placement to be confirmed before implementation (no inline controls on HMI cards). Alarm + AuditRecord events cover fence, re-bind, and lease loss (21 CFR Part 11). - The running phase survives a failover. Phase/SFC state — down to
active steps, variable values, and fired transitions — lives in
Phase.Statusin etcd, not on the edge node; the batch record shows the exception (Hold, Restart) rather than a vanished phase. - Hardware guidance changes: reference architectures stop calling
the edge node an unmitigated single point of failure and instead
document the failover-target patterns — dedicated standby per
controller (recommended) vs. shared spares (cost option,
co-location trade-off).
docs/ha-failure-modes.md,docs/reference-architectures.md(Pattern B), anddocs/production-deployment.mdupdate when the mechanism ships; IEC 62443 FR 7 availability targets for the unit runtime can rise accordingly. - Marketing posture: until this ships, the accurate claim remains autonomy + restart-replay + supervised Hold. Once shipped, the claim is controller failover (hold-then-resume) to designated standby hardware — never "bumpless redundancy" or redundant-pair parity with DeltaV.
- Default behavior is unchanged:
Autonomymode preserves today's semantics exactly; everything else is additive and opt-in. - Followups: the implementation epic (#574). Warm-standby pre-staging (image pre-pull, a pre-created fenced pod on the designated standby) is a compatible later optimization that shortens failover time without state replication.
- Reversibility: high before implementation (this is a posture); moderate after — the availability policy field becomes a public CRD contract, and withdrawing automatic failover would be a customer-visible regression requiring a superseding ADR.
Amendment (2026-08-07, #1320): the automatic path issues the Hold itself¶
Hold-then-resume as implemented waited, before re-binding a Running unit, for "the crash detector or watchdog" to issue the Hold. Both of those live in the pod-lifecycle pass of the same reconcile, downstream of the failover check — so the wait starved the only pass that could end it, and an automatic failover of a unit Running a batch parked forever (found live: eight minutes on a dead node with no Hold, no alarm, no re-bind).
The decision is unchanged; the sequencing within it is corrected. A
lease observed Expired means the runtime has self-fenced, which is
the crash evidence itself, so the automatic path now issues the ISA-88
Hold directly — through the same code path, alarm, and
RuntimeCrashDetected condition as the crash detector — once the
expiry outlasts the ADR 0008 hold bound (½ lease, clamped 5–30s). Inside
the bound it waits, so a runtime that recovers and renews its lease
rides through with no Hold, preserving the crash detector's grace
semantics with a lease-sized budget. The re-bind still waits for the
safety margin (margin ≥ hold bound by construction), and the lease
expiry is still reported (status, alarm, audit record) before any Hold
or re-bind. The manual path's refusal of Running units is unchanged.
Amendment (2026-08-31, #1890): start-up is in the lease arithmetic, and the automatic path spends a budget¶
One control-plane node loss on the bench produced 31 failovers of a controller that never failed, over 23 minutes, and a 1924.6s hole in the process record. It stopped on its own and nothing was done to stop it. ipc-2 was up throughout, and the node that went down carried no runtime and no field path. The same shape had fired three times before, unreported.
A re-bind destroys the thing whose liveness decides whether to re-bind. The newly-bound runtime has to start, pull its image, load its program, connect its drivers and get one renewal acknowledged, and the only window it had for all of that was the lease duration. A lease duration is a renewal budget: it says how long an established runtime may go silent before it has certainly self-fenced, and it is short because it is also how long a dead controller keeps the plant waiting. Spending it on start-up made the mechanism self-sustaining under exactly the conditions that trigger it — a degraded control plane makes the registry, the broker and the field all slow at once. Nothing damped the loop: no backoff, no cap, and no grace for a runtime that had been bound seconds earlier.
It was invisible because every symptom reads as something else. The alarms are
failover and lease-expired, which are the correct alarms for a controller
that genuinely died. The unit settles Idle on a healthy node, which is a
correct-looking resting state. docs/ha-failure-modes.md predicts a failover
on lease loss, so the first one matches the documentation exactly. Only the
count and the epoch give it away, and nothing surfaced either.
The decision is unchanged. Two things are added to it.
1. The lease has two deadlines, chosen by whether the runtime has ever
answered. For a binding this operator has had a renewal acknowledged at, the
deadline is unchanged and must stay unchanged: lastSuccess + leaseDuration,
which is what the whole one-writer ordering above is built on. For a binding
where no renewal has ever been acknowledged, the deadline is
availability.runtimeStartupGraceSeconds (default 120s, never shorter than the
lease) from the instant the runtime pod first had an address. Expired on that
second path is not a detection of a self-fence — no lease was ever granted at
that epoch, so nothing is driving anything and there is nobody to fence. It is
a verdict that the replacement failed to come up, and it is sized accordingly.
Delaying that verdict can only delay a re-bind, never advance one, so nothing about the one-writer guarantee moves. The asymmetry is what makes a generous default right: too long delays a re-bind for a unit that has no runtime either way, and too short destroys the pod that was seconds from answering.
2. The automatic path may try each free eligible standby once between
confirmed runtimes. status.runtimeBinding.unsettledRebinds names the nodes
an automatic re-bind has already been sent to since the runtime was last
confirmed operating normally. A target in that list has been tried and did not
establish, so choosing it again repeats an experiment whose result is in hand.
When every free eligible target is named there, the automatic path suspends:
it sets FailoverSuspended, raises a Critical alarm of its own — deliberately
not under the ordinary failover alarm's prefix, because reading as an ordinary
failover is how this hid — and re-binds nothing.
Suspension is a stop on re-binding and not on recovery, and that distinction is
what makes it safe. The current binding's pod stays where it is, the broker
goes on renewing against it, and the first renewal it acknowledges takes the
lease to Held, which clears the condition, the alarm and the ledger through
the same edge that clears an ordinary failover's alarm (#1733). In the bench
episode that alone would have been the recovery: what prevented it was the
re-bind destroying, every 45 seconds, a pod that was 25 seconds into starting.
A deliberate dcs unit failover remains available throughout and empties the
ledger, because a human choosing a target is a new decision rather than a
repetition of this controller's.
A timed automatic retry was considered and refused. It re-enters the same loop at a longer period, and what it retries has already been shown not to work.
3. The expiry says which of the two findings it is. The arithmetic above
splits Expired into two events, and for a while only the arithmetic knew.
The Info log, the Critical alarm and the Part 11 audit record all went on
saying that the runtime had self-fenced and stopped writing. On the
never-established path that describes something which did not happen, and it
is the sentence an operator reads while deciding whether a controller died.
Thirty-one of the bench episode's thirty-two alarms said it about a
replacement pod that was still starting.
The broker now captures its verdict at the instant it declares the lease
Expired, and all three records read it. An established lease keeps the
sentence it had. A lease that was never granted is annunciated as a start-up
that failed, naming the grace it was given. The verdict is captured on the
transition rather than read afterwards, because the fact that separates the
two is whether any renewal has ever been acknowledged, and that is precisely
what the runtime finally answering erases.
The verdict also carries what the last renewal attempt got back, and whether
the runtime's own handler answered it. A refusal and a lost connection are
opposite findings about the same failed renewal
(ADR 0081): the refusing runtime
is up, reachable and talking, while an unreachable one may be a dead node.
Both used to end at V(1), which ADR 0063 makes invisible in a shipped
binary. That is why the bench episode's opening expiry could not be explained
when this was written: sixty-odd renewals failed against a runtime that was
believed to be up, and no durable record anywhere said what a single one of
them returned. Part 5 below has the answer, read back off Prometheus.
The audit reason is unchanged on both paths, and deliberately so. The batch-record assembler matches on it to bound the failover data gap (#1308), and a replacement that never came up stops the samples just as surely as a runtime that fenced. What the two paths owe the reader is different prose, not a different correlation key.
4. Whether a runtime has ever answered is a fact about the binding, not
about the operator process asking. The two paths above both turn on one
question, and the broker answered it out of its own memory. That memory lives
in the operator process and dies with it, while the runtime it describes does
not: a replacement physical-operator starts a renewal loop against a runtime
that may have held the lease for a week and sees exactly what it would see at
a pod that has never answered anybody.
Told nothing, the replacement applies both halves of this amendment to an established runtime. It waits out a start-up that finished days ago, which delays the expiry the re-bind margin runs from by the whole of the grace, and then annunciates that expiry as a replacement that failed to come up, about a controller that really did self-fence. Neither is a finding about the plant. Both are the new process's ignorance rendered as one.
status.runtimeBinding.leaseEstablishedBy records it instead. It names the
runtime incarnation — pod UID and the runtime container's restart count —
that has had a renewal acknowledged at this binding, written once, on the
first acknowledgement, by whichever operator observes it. A binding with
nothing recorded is one no runtime has ever answered at, which is what a fresh
binding looks like and what every re-bind produces, since the re-bind builds
the object fresh.
The incarnation and not the pod is what is recorded, for the same reason
TerminalStopArming carries the pair. A container that restarts in place
keeps its pod, its UID and its binding, and loses every piece of state the
lease was established with. That process is starting up and has to earn the
lease again, which is precisely the case the start-up budget exists for.
A pass with no acknowledged contact behind it writes nothing. Recording an incarnation this operator has not heard from would hand the next one a claim nobody ever had, and would spend the start-up budget of every fresh binding — which is the storm.
This is the bench episode's own opening loop. A replacement operator took leadership 372 seconds after the node loss and began renewing against ipc-2, bound at epoch 60 throughout, and 30 seconds later declared the lease expired. Why those renewals were not acknowledged is still unanswered, and this changes nothing about that. What it changes is that the answer, when the next occurrence gives one, is no longer read against a runtime the control plane has mistaken for a new one.
5. What the bench episode's two unexplained halves were. Read on 2026-09-03 out of the bench's Prometheus, whose retention reached back past the storm, against T0 = the cord out of cp-1 at 23:50:50Z on 2026-08-31.
The opening expiry was #1918. The runtime on ipc-2 fenced at T0+15 s on its
own clock, with the grantor gone. On chart 0.6.3 the liveness probe answered
from the same verdict as readiness, so a fenced runtime was NOT_SERVING to
both, and the kubelet killed it about a minute into the fence. A Failover
runtime starts fenced, so the restarted container was NOT_SERVING from its
first probe and was killed again. The restart counter read 1, 2, 3, 4 and 5 at
T0+80 s, +140 s, +200 s, +260 s and +320 s, and kube-state-metrics reported the
container in CrashLoopBackOff from T0+380 s to the re-bind at T0+422 s. The
replacement operator's broker renewed from T0+377 s. Every one of its renewals
was a TCP connect refused by a node whose container the kubelet was holding in
backoff, which is the transport class of ADR 0081. "ipc-2 was up throughout"
was true of the node and false of the process.
The 31 cycles after it were #1933. The broker's persistent volume is pinned to
cp-1, so its replacement pod sat Pending from T0+380 s until cp-1 returned at
T0+1760 s. A fresh unit-runtime blocks inside Start on the MQTT connect,
with no timeout, and opens neither its health server nor its lease endpoint
until the broker answers. Each of the 31 pods was created, blocked there, and
was deleted by the next re-bind 45 s later. None reached Ready and none
restarted. The pod created at T0+1788 s, 28 s after the broker scheduled, is
the one that took the lease.
Neither half changes this amendment. The start-up grace is still the right deadline for a runtime that has never answered, and the budget is still what stops the loop. On that night it would have stopped after two re-binds instead of 31, and the pod it left in place would have taken the lease the moment the broker did. What the two halves add is the two causes of a suspension that this ADR's runbooks did not list: a kubelet restarting the runtime the grantor is waiting for, and a broker the replacement cannot reach.
Amendment (2026-09-01, #1894): the grantor keeps granting when it loses the apiserver, and the operator that replaces it waits¶
Cutting one control-plane node on the bench fenced a unit on a node nobody touched. The node that was cut held the Talos L2 API VIP, the unit's runtime was on an edge node, and the operators were on a third node that stayed up throughout. Nothing the unit depends on was in the blast radius. It self-fenced at T0+29.8s, its lease duration to within the sample interval, drove both outputs to zero and stayed fenced for 37 seconds.
The chain is short and every link is behaving as designed. The API at the VIP
was unreachable for 57.615s while the address floated to another node. Every
operator pod reaches the apiserver through the in-cluster Service, and every
operator whose connection was pinned to the lost endpoint timed out for the
whole of that window (the bench verification below is what put the word
"pinned" in that sentence). controller-runtime
terminates the manager when leader-election renewal fails, which is the correct
behaviour for a reconciler: a controller that cannot reach the apiserver cannot
be trusted to be the only writer of the objects it owns. The lease broker lives
inside physical-operator, so the process exiting took the plant's grantor with
it, and the 30s lease expired under a runtime that never moved.
It is the wrong behaviour for a liveness grantor, and the reason is the whole amendment. A grant of liveness is carried on a direct HTTP channel to the runtime's pod IP and never touches the apiserver. Losing the apiserver removes this process's power to move a unit. It does not remove its ability to keep one alive, and it does not remove the duty. The two were tied together only because they happened to share a process.
This is not the failure #1893 answers and the two must not be read as one. There the node carrying the operator is lost, the clock in front of it is the 300s pod-eviction timeout, and a second replica is the answer. Here the apiserver is lost, the clock is a VIP failover measured at 57.6s, and a second replica helps only by chance. The Service address is shared, but each pod's connection is pinned to one apiserver endpoint, so the follower rides the outage through when it is pinned elsewhere and dies with the leader when it is not. Acquiring the leader lease is itself an API write. A graceful drain of the same node removes the API outage entirely and hands every lease off within ±3s, and the unit still expires a lease and still fences, for 6.4s. Only the duration changes.
Three things are added to the decision.
1. A physical-operator whose manager stops without being asked to stands down
instead of exiting. For physicalOperator.leaseStandDown (default 90s) both
edge liveness brokers go on renewing what they already hold. They learn no new
targets, re-bind nothing and write nothing, because the reconcilers that would
do any of that stopped with the manager. Then every loop is cancelled and the
process exits. A runtime whose control plane came back never notices; a runtime
whose grantor never returns fences at the end of the window, which is the
outcome the fence exists for. The window is sized to outlast the platform event
that takes the apiserver away and then gives it back. That event has a ceiling:
a Talos L2 VIP is held through an etcd concurrency session at etcd's default
60s TTL, so an abrupt loss of the holder costs at most what is left of that
lease (#1905), and every healthy-quorum figure the bench has produced sits just
under it. 90s clears the ceiling by 30s. The first draft of this amendment
called it "the measured 57.6s with half of itself again", which was one sample
and headroom. A shutdown signal ends it early,
because a pod being deleted deliberately is not a pod that lost the API.
The process keeps answering /healthz while it stands down, which is a
mechanism detail with a safety consequence. The manager serves that endpoint and
has just stopped serving it, so without something answering, the kubelet's
liveness probe kills the container most of the way through the very window it
was given.
2. An operator that has just taken leadership does not read a lease expiry as
a fence. This is the price of the stand-down and it is paid by whichever
process takes over. The dangerous shape is a partial partition, where the old
leader keeps reaching the runtime while the new leader cannot: the new leader
watches its own renewals fail, calls the lease Expired, and re-binds onto a
standby while the old runtime is still being held alive. Two writers, which is
the one thing this ADR exists to prevent.
The arithmetic that closes it uses only quantities both sides know. The old
leader gives up leadership at some instant T, and a challenger cannot acquire
before T, because leader election hands the lease over strictly after the
holder's renew deadline has passed. The old leader stops renewing at
T + leaseStandDown, and the runtime it was holding fences at most one lease
duration after that, plus the bounded ADR 0008 pre-fence hold the re-bind margin
already covers. So a new leader that waits leaseStandDown + lease duration +
margin from its own acquisition has waited past that fence, whatever the
partition did to the two of them.
The hold-off gates every re-bind whose fence evidence is the lease expiry, on
the automatic road and the manual one alike. It does not gate a re-bind a human
has certified with dcs.io/failover-confirm-fenced: powering a node off is a
stronger fence than any arithmetic here can construct, and it is asserted about
the plant rather than about a clock. The wait is written onto the unit as a
FailoverRequest condition with reason LeadershipHoldOff, because a wait
nothing renders is the defect #1198 fixed. The condition names the deadline
rather than the remainder, and it is the one reason on that condition the
operator removes on its own, on the first pass after the deadline with no
re-bind in between (#2148). Every other reason is an outcome and stands until
the next outcome. This one is a wait, and the bench read "blocked for a
further 1m29s" ten days after the wait had ended.
Both halves travel on one value. Setting physicalOperator.leaseStandDown to
zero disables the stand-down and the hold-off together, which is the pre-#1894
behaviour: losing the apiserver fences every Failover unit at its lease
duration and latches every Autonomy one.
3. A physical-operator that is stopped on purpose hands its leadership straight back. The stand-down answers the abrupt cut and does nothing at all for the graceful one, because a pod being deleted deliberately is not a pod that lost the apiserver, and the two must not be confused. The graceful drain nonetheless fenced a unit for 6.4 seconds with the apiserver never away and every operator lease handed over within ±3s, so something in that path was still too slow, and it was the handover itself.
LeaderElectionReleaseOnCancel was off, which is the controller-runtime
scaffold's default. A leader shut down on purpose therefore kept its lease until
the full LeaseDuration ran out, so a standby could not begin syncing its
caches for fifteen seconds and only then did its broker start renewing unit
leases. Against a 30s unit lease that arithmetic has almost nothing left in it,
and on the bench it had nothing.
It is now on, for this operator and not for the other four. This is the only one whose absence reaches the field, so it is the only one whose leader-handoff latency is charged against a plant-side budget rather than against reconciliation lag. The scaffold leaves the option off with a warning, and the warning is exactly the thing this ADR now has to keep true: it requires the binary to end immediately when the manager is stopped. Since point 1 this binary does not always end immediately. The two are compatible only because they are disjoint on precisely the condition the warning is about — a release happens when the manager was cancelled, and the stand-down runs only when it was not — so the branch and the option have to be read together and changed together.
Read as one sentence: the authority is handed over as fast as possible when it is being given up on purpose, and held for as long as is safe when this process is only out of touch.
Bench verification (2026-09-02, chart 0.7.1) and two corrections¶
Three drill 12 reps on the first release carrying both commits, every one an abrupt cord pull at the PDU. Rep 2 is the verification of point 1: the leader, pinned to the cut node's apiserver, lost leadership at T0+14s and stood down; no lease expired and no unit fenced; and the incoming leader was renewing the unit's lease six seconds before the standing-down one released it, so the handover overlapped rather than gapped. Point 3 and the plant cost were measured the same evening, once the bench had a cut phase long enough for the cord to come out with the loop live (every earlier rep that triggered the fix had its field half void, because the demonstration phase's 120s dwell left a 20s cue window). A graceful drain of the node carrying the leader and the VIP moved the grantor lease at T0+1s, expired nothing and fenced nothing, AO held and PV band 59.998 throughout: point 3 verified. An abrupt cut of the node carrying the VIP and the leader's pinned apiserver, leader on a third node and standby pinned elsewhere, engaged the stand-down at T0+15s, had the standby renewing at T0+24s, the API back at T0+59.0s, and the old process releasing at the end of its window at T0+105s with nothing left to regain; no expiry, no fence, AO held at 19628, PV in band. The cost of the fault to the plant was zero on both reps.
First correction: the discriminator is the pinned endpoint and not the VIP.
Rep 1 cut the VIP holder and verified nothing, because the leader never lost
its API. On the same surviving node the control-operator hit the exact
signature at T0+9s and exited while the physical-operator beside it never
noticed. In-cluster pods do not use the VIP. The Service address is DNAT'd to a
node address, and a pod's long-lived connection stays pinned to that endpoint
until the connection dies. So an abrupt control-plane loss reaches roughly one
pod in three, whichever node holds the VIP, and the amendment's claim that a
second replica "is no answer at all" was too strong. It is an answer two times
in three, and the stand-down is the answer the third time. A rep that wants to
exercise the stand-down reads the leader's conntrack entry first and cuts the
node it names.
Second correction: the window is sized against a ceiling. Rep 3 cut the node holding the VIP and the leader's pinned endpoint together, the stand-down engaged at T0+15s and released at T0+105s as designed, and the unit fenced anyway, for 165s, because the API was away for 307s. That reading is not evidence about the window. cp-1's etcd had been dead since rep 2 with a corrupted store, so cutting cp-2 was the second fault of two and took quorum with it, and without quorum the VIP cannot move at all (#1905). A VIP failover on a healthy quorum is bounded at 60s by the etcd session TTL that elects the holder, and 90s clears it. A quorum loss is bounded by nothing, and a plant that would rather ride one through with no failover capability raises the chart value.
Why the window stays a timer. Rep 3 raised the question of whether the
stand-down should end on a condition instead, since the grantor let go with a
working channel to the runtime still in hand. It cannot. Point 2's arithmetic
is finite only because point 1's window is: the hold-off exists for the
partition where the old grantor reaches the runtime and the new leader does
not, and in that partition neither "another operator has taken over" nor "I can
no longer reach this runtime" ever arrives at the old grantor. With no bound on
its tenure there is no instant at which the new leader may read an expiry as a
fence, and a Failover unit whose runtime really has died is then never
re-bound on the automatic road. The timer is the bound, and the chart value is
where a site chooses it.
Amendment (2026-09-02, #1909): the grantor gap is sized against the budget after a renewal¶
The #1893 second replica was bench-read the same day it shipped, and the
arithmetic it was written against did not hold. Drill 12 rep c cut the node
carrying the physical-operator leader, with the standby pinned to its own
apiserver so nothing but the grantor's death was in the blast radius. Leader
election did what was predicted: the standby acquired at T0+16.7s against a
15s election lease polled every 2s, and the plant's spare grantor was back at
T0+101s against 397s for the single-replica procedural-operator on the same
cut. The unit fenced anyway, for 9.5s, from T0+11.9s to T0+21.0s, with the
batch idle so the field cost went unmeasured.
Why the fence landed at +11.9s and not at +30s. The 2026-08-24 amendment
above anchors the runtime on the renewal the operator had acknowledged, which
is the one BEFORE the POST that carries the report. After a renewal at t the
runtime's deadline is therefore t minus the renewal interval plus the lease,
and the budget a replacement grantor has to fit inside is the lease minus the
interval, less another interval depending on where in the cadence the cut
landed. At a third of the lease that was 10 to 20s of a 30s lease. On the rep
the last renewal was at T0-8.5s and the fence at T0+11.5s, to the sample. The
standby needed 16.7s to acquire and 4.3s more to its first renewal, which is
the Unit reconciler reaching reconcileAvailability through everything a Unit
reconcile does first. So a two-replica grantor fenced the unit on EVERY abrupt
loss of the leader's node, for between one and ten seconds.
The anchor is not the defect. It is what keeps the operator's Expired
decision and the runtime's self-fence on one instant, and it is why #1890's
re-bind cannot land on a runtime that is still driving. Moving the runtime's
anchor to the arrival of the POST would put the fence ten seconds after the
operator's own expiry and reopen the two-writer hazard this ADR exists to
close. The grantor gap is what has to shrink, and the budget it has to fit
inside can be widened without touching the anchor. Three changes, and the
decision is that all three ship together, because no one of them alone
closes it:
-
The election is this operator's own. controller-runtime's 15s/10s/2s is a scaffold default for a reconciler, where a lost election is a restart and nothing in the field is timed against it. On the
physical-operatorthe election is the whole of the grantor gap. It runs at 6s/4s/1s (LeaderElectionLeaseDurationand its two siblings, carried by the--leader-election-*flags andphysicalOperator.leaderElectionin the chart), which puts acquisition at about 8s after an abrupt loss. The 2026-09-01 amendment is what makes the shorter renew deadline affordable: a leader that loses a renewal stands down and goes on granting what it holds, so a spurious loss costs one leader change and nothing in the plant. The other four operators keep the scaffold default. -
Every liveness loop starts from the cache on acquisition. A leader-election Runnable (
LivenessWarmStart) lists the Units the cache already holds the instant this process wins, and starts the lease and heartbeat loops for every runtime pod the reconciler would have started them for. It builds each target with the same two functions the reconciler calls (leaseTargetFor,heartbeatTargetFor), so the two roads cannot decide differently about one pod, and it writes nothing. The 4.3s becomes one cache list. It is behind the election on purpose: a follower granting liveness would be granting to runtimes it has no authority to move. -
The renewal cadence is a sixth of the lease.
renewIntervalwasduration/3, which made the budget after a renewal two thirds of the lease. Atduration/6it is five sixths, 25s of a 30s lease, and 20s in the worst phase. The cost is one POST per unit every 5s instead of every 10s. The #1810 clamp and its tests are unchanged, because the rule they state is about the anchor and not the divisor.
Together: a worst gap of about 10.4s (6s election, two jittered 1s polls, 2s
budgeted for the warm start) against a worst budget of 20s.
TestGrantorGapFitsInsideTheBudgetAfterARenewal holds the three constants
against each other with 5s to spare and fails if any one of them reverts, and
deploy/helm/cloud-native-dcs/tests/leader-election.sh holds the chart's
rendered values against the same inequality, because a values.yaml edit
reaches no Go test.
What this amendment does not claim. The fix is unverified on the bench.
The rep that proves it is the same rep c, aimed the same way, and it reads
three things the first rep could not: the standby's acquired to its first
starting lease renewal loop, which is the warm start's real figure and
replaces the 2s budgeted above; the runtime's own control lease expired line
minus its leaseDuration, which is the anchor the fence ran from and the only
record of the last renewal once the dead leader's log has gone with its node;
and whether the unit fenced at all. A fence of any length on that rep reopens
this amendment.
A site whose apiserver cannot answer a Lease update inside four seconds widens all three election timings together, and pays for it in the budget above. client-go refuses a set where the lease does not exceed the renew deadline, or the renew deadline does not exceed the retry period with its jitter, and it refuses at manager construction, so a misconfigured operator exits before it can grant anything.
Amendment (2026-09-02, #1918): a fence is not a death, so the two probes ask two questions¶
Drill 12 rep 3 of the #1894 verification fenced a unit for 165s. That was a quorum loss (#1905) and the fence itself was correct. Inside it the runtime container exited and was restarted by the kubelet:
19:01:38.253 control lease expired; running bounded edge-local hold before self-fencing
19:01:38.362 control lease expired, self-fenced: FB output writes stopped, reads continue
19:02:20.417 shutting down unit runtime
(exit 1; container restarted 19:03:51)
The pod's liveness and readiness probes both asked the same gRPC health
service on 61052, and the runtime answered both with one verdict. A fenced
runtime answered NOT_SERVING, which is right for readiness and was fatal for
liveness: three failures at a 20s period landed the restart about a minute into
the fence. Every fence shorter than that was unaffected, which is why #1909's
9.5s and the original 37s never showed it.
A fence is the runtime doing the one thing this ADR asks of it. It is not
an unhealthy process. Restarting it inside the fence throws away the process
that ran the hold program, forces the self-held snapshot through the #1804
recovery road, and on the bench opened the historian gap at the same second
(Data gap opens vs T0=+165.950s against a restart at T0+166s). The 165s fence
in that rep was two events, a fence and a restart, and the row recorded one.
The #1894 stand-down exists so a runtime survives a control-plane absence
untouched, and this was the kubelet touching it on a schedule.
The same verdict had a second victim that the bench did not have to show.
Every field driver gone past the 60s grace was also NOT_SERVING, so the
liveness probe restarted a runtime for losing its plant. A restart reconnects
no cable. The reconnecting driver is the in-process remedy, and the restart
threw away the held state while it ran.
The decision. Readiness and liveness are different questions, and the
gRPC health protocol can answer several named services from one server, so
they are two services (pkg/grpc: ServiceReadiness, ServiceLiveness).
- Readiness (
readiness, and the protocol's unnamed overall status) is what it has always been:NOT_SERVINGwhile fenced, andNOT_SERVINGwhen every field driver has been gone past the grace period. The unit controller reads pod readiness to see the outage, and that path is unchanged. - Liveness (
liveness) answers whether the process is alive and scanning. Fenced, held and degraded are allSERVING. The oneNOT_SERVINGit can give is a network that reports itself Running or Degraded and has not completed a scan in longer than its stall budget, 60s or ten scan intervals, whichever is longer. A wedged scan goroutine holds the plant's last commanded state and regulates nothing, nothing inside the process can free it, and a restart is the remedy. That is the only case where it is.
The health server refuses the liveness service with NotFound when it was
built without a liveness verdict, and refuses any name it does not know. A
probe naming the wrong service therefore fails loudly on the first probe, never
quietly on the readiness verdict under another name.
What holds it. The unit runtime's pod is built by the physical-operator
and not by the chart, so the gate is
TestUnitPodProbesNameDifferentHealthServices in
internal/controller/physical: the liveness probe names ServiceLiveness, the
readiness probe names ServiceReadiness, and the two constants differ. It was
proved by mutation: pointing both probes at the readiness service fails it on
the liveness line. TestLivenessCheck_FencedRuntimeIsNotReadyAndAlive holds
the two verdicts apart on a fence, and TestLivenessCheck_StalledScanIsDeath
holds the one case where liveness does fail.
What this amendment does not claim. The restart was read out of one
container log, and the fix is unverified on the bench. The rep that proves it
is any fence longer than about 90s with the runtime's restart count read before
and after. The row in ha-failure-modes.md that said the kubelet restarts the
pod after all drivers are lost was bench-validated on the Hold at +21.1s and
not on the restart, and it no longer says so.
Amendment (2026-09-03, #1930): a dead apiserver connection is found by the ping, and the ping is timed against the election¶
The #1909 amendment was bench-read the day after it shipped, on chart 0.7.3,
drill 12 rep c: cp-3 cut at the PDU with the physical-operator leader and
the VIP on it, the standby on cp-2. The three fixes did what they shipped to
do. From the moment the standby could reach an apiserver it acquired within
the 6s election lease, and acquisition to starting lease renewal loop was
1.4s against the 2s budgeted. The unit fenced anyway, for 29s, from T0+23s to
T0+52s, because for 43s the standby had no apiserver at all. Twelve lease
reads in a row timed out at 2s each, and then at T0+43s an unrelated audit
write logged http2: client connection lost, the standby dialled again, and
the election ran.
The 43s was client-go's, not the node controller's. The issue as filed
read the standby's recovery as the kubernetes Endpoints being pruned when
cp-3 went NodeNotReady at T0+44s. That was a coincidence of two clocks. The
apiservers prune each other by lease, 15s TTL reconciled every 10s, and node
status has no hand in it. What held the standby was its own connection. Every
in-cluster client dials the Service address, kube-proxy DNATs the connection
to one apiserver endpoint, and client-go keeps that HTTP/2 connection for as
long as it looks alive. A node losing power sends no FIN and no RST, so the
connection looked alive, and every request rode it until the transport's
health check gave up: a PING after 30s with no frame received, a close after
15s more with no answer. The standby's last frame from cp-3 arrived a moment
before the cut, and 43s later the connection was declared lost. #1894 had
found the same pin deciding the LEADER's fate and answered it with the
stand-down. The standby has nothing to stand down from. It can only wait for
its transport, and its transport was waiting on a 45s clock nobody had put in
the arithmetic.
The decision. The two timeouts are client-go's own, read from the process
environment (HTTP2_READ_IDLE_TIMEOUT_SECONDS, HTTP2_PING_TIMEOUT_SECONDS)
when a transport is built and from nowhere else, so the physical-operator
sets them before it builds one: 2s and 2s (APIServerReadIdleTimeout,
APIServerPingTimeout, flags --apiserver-read-idle-timeout and
--apiserver-ping-timeout, chart physicalOperator.apiserverHealthCheck).
Detection inside four seconds of the last frame, in place of forty-five.
The detection does not overlap the election lease, and the arithmetic has to
say so. A challenger restarts its lease clock the first time it reads a record
it has not seen, because it cannot know how long ago the holder wrote it. The
leader renews every second and the standby reads every second, so a renewal
between the standby's last read and the cut is the usual case, and the first
successful read after the reconnect is that unseen record. The worst gap is
therefore the detection, then a poll, then the whole election lease, then a
poll, then the warm start: 4 + 6 + 2.4 + 2, about 14.4s, against the 20s the
runtime has after a renewal in the worst phase, with 5s to spare.
TestGrantorGapFitsInsideTheBudgetAfterARenewal carries the fourth term now
and fails at client-go's defaults, where the sum is 55s.
deploy/helm/cloud-native-dcs/tests/leader-election.sh holds the chart's
rendered values against the same sum, refuses a read-idle of zero (client-go's
documented disable, which restores the 43s) and a fraction of a second (which
client-go would read as something else), and clears the block to prove the
flags are the values and not the template.
TestAPIServerHealthCheck_ADeadConnectionIsGivenUpInsideTheWindow is the
mechanism itself, driven through rest.HTTPClientFor against a relay that
goes silent without closing, which is what a node losing power looks like to
its peer. The tuned client re-dialled 2.02s after the peer went silent at 1s
and 1s. The control case, the same fault under client-go's defaults, kept the
dead connection for the whole window, which is the bench's 43s reproduced and
is what proves the environment moved it rather than the relay.
The price. One PING per connection after two quiet seconds, and a fresh
TCP and TLS handshake when an apiserver cannot answer a PING inside two more,
with the process's watches resuming on the new connection. On the leader a
spurious close can cost a renewal, which since the 2026-09-01 amendment costs
one leader change and nothing in the plant. What it can never cost is a grant
of liveness, which travels to a runtime pod IP and never touches the
apiserver. The other four operators keep client-go's defaults: a reconciler
that sits on a dead connection for 45s restarts on leader election lost,
and nothing in the field is timed against it.
What this amendment does not cover. A fresh dial while the dead endpoint
is still in the kubernetes Endpoints. For up to 25s after the cut, until the
apiservers' lease reconciler drops it, a new dial lands on the dead node one
time in three on a three-node control plane and hangs for the election
client's 2s request timeout before the next poll dials again. Each such dial
is one term of the 5s margin, so two in a row spend it in the worst phase and
three fence the unit, which is one rep in twenty-seven at the worst phase and
fewer elsewhere. The issue's alternative closes that residue too, at a cost
this amendment does not take: dialling the node-local apiserver over
status.hostIP gives the standby a connection that cannot die with the
leader's node, and it requires every physical-operator pod to be scheduled
on a control-plane node, which the reference bench does and a production
install need not. That stays an opt-in for the founder to rule on, not a
default.
What this amendment does not claim. The fix is unverified on the bench.
The rep that proves it is rep c with the standby's pin READ as the leader's
node before the cut, which the census could not do at the cue because a
standby's 2s election reads slip between conntrack samples. The reading is
the standby's own log: its last Error retrieving lease lock to
Successfully acquired lease should be under 10s, the runtime should log no
control lease expired, and dcs_runtime_fenced should stay at 0. A fence of
any length on that rep reopens this amendment. Rep c as the #1909 amendment
specifies it, standby pinned to a surviving apiserver, is a different rep and
is still owed.
Amendment (2026-09-03, #1933): a broker is a transport, and a runtime starts without one¶
The 2026-08-31 storm had two halves the #1890 amendment did not explain, and
both were read back off the bench's Prometheus on 2026-09-03. kube-state-metrics
was scraped throughout, with a week of retention. The opening expiry was #1918:
five liveness kills of a fenced runtime, then CrashLoopBackOff across the
replacement grantor's renewal window. The 31 cycles that followed were this
defect.
What the runtime did. Adapter.Start connected the device drivers
non-fatally and then called the MQTT client's blocking connect with the process
context, which is cancelled on SIGTERM and by nothing else. autopaho's
AwaitConnection returns when a connection is up or the context is done.
cmd/unit-runtime opens the gRPC health server and the HTTP API only after
Start returns. So for as long as its broker was unreachable a fresh runtime
listened on nothing: no readiness, no liveness, and no lease endpoint. A
Failover runtime could not be granted its control lease, and it could not say
so. An established runtime was never affected. Its client reconnects in the
background, the store-and-forward queue holds its publishes, and its lease
keeps renewing. The defect was on start-up only, which is exactly the phase
the #1890 arithmetic measures.
What it did on the bench. The broker's PersistentVolumeClaim is
local-path, so it pinned the broker to cp-1, the node that was cut. The
replacement broker pod was created at T0+380s and stayed Pending until cp-1
came back at T0+1760s. Between T0+422s and T0+1788s the automatic path created
31 runtime pods, one every 45.5s, alternating between the two device nodes.
None reached Ready and none restarted. Each was a container blocked in Start,
deleted by the next re-bind. The pod created at T0+1788s, 28s after the broker
was scheduled, took the lease at T0+1805s. The storm ended when the broker did,
and nothing else changed at that minute. The healthy re-bind on 2026-08-27 had
taken 10s to Held with the broker up.
The decision. A broker is a telemetry and audit transport. It is not a
precondition of control, and the start path now says so the way the hold path
has since #1779. Start opens the connection with the client's
ConnectBackground and returns. The error that call can return is a
configuration error, an unparseable URL or unreadable TLS material, and never
reachability. The retained status a starting runtime publishes rides the
store-and-forward queue while the broker is away, in order ahead of anything
published later, and the replay worker's on-connect listener delivers it when
the broker answers. The health server and the lease endpoint therefore come up
on a node whose broker is away. Readiness still answers to the lease, as it did.
A returning broker is also told what the runtime is now. The queue can evict the start-time status in a long outage, and the broker that comes back can be a fresh one holding no retained state. So every connect re-announces the current status, and the announcement goes through the queue rather than around it. A live publish from the connect listener would race the replay of an older status still queued, and the older one landing second would leave the broker asserting a fence state the runtime had already left. The queue is FIFO, so the announcement is delivered after everything that was true before it.
The mqttPublisher slice the runtime holds carries ConnectBackground and not
Connect. The blocking form is the one call in the client that can hold a
goroutine for the length of an outage, and the runtime has no goroutine it can
spend that way. Leaving it off the interface is what keeps it from coming back.
The tests. TestStart_ALeaseIsGrantedWhileTheBrokerIsUnreachable starts a
Failover runtime against a closed loopback port, bounds Start at ten seconds,
grants the lease over the HTTP endpoint with the client still reporting
disconnected, and finds both status publishes in the queue. With the blocking
connect restored it fails at the bound, so the bound is the assertion.
TestStart_TheStatusIsAnnouncedOnEveryConnect brings the fake transport up,
grants the lease, takes the transport down and brings it up again, and reads
the last status on the topic each time. Without the connect listener it fails
on the first connect.
What this amendment does not cover. The device drivers are still connected
inside Start, non-fatally, and a driver whose dial has a long timeout still
delays the servers by that timeout. Nothing on the bench has measured that
phase. (The same bench reading measured it that afternoon, and the #1936
amendment below answers it.) A runtime with the queue disabled (--mqtt-queue-size 0) publishes its
announcement live on a goroutine of its own, because there is nothing for it to
race, and it drops the start-time status the way it drops everything else the
broker did not take. The bench rep that found this, drills/rebind-budget-rep.sh
in cndcs-deploy-bench, reads the replacement Held inside the start-up grace
with the broker still down once a chart carrying this amendment is on the rig,
and its re-bind budget claims do not apply on that chart. The bench is pinned to
0.7.4, which predates this, so that reading is owed.
Amendment (2026-09-03, #1936): a driver is dialled beside the start, and a dial that failed is still watched¶
The #1933 amendment closed with the phase it did not cover: the device drivers
were still dialled inside Start, one after another, and nothing had measured
it. The bench reading that closed #1935 the same afternoon had the numbers in
the replacement pod's own log. starting runtime at 17:16:33.215, failed to
connect IO driver, will retry on first I/O at 17:16:36.281 with dial tcp
10.10.20.51:502: connect: no route to host, and runtime started successfully
three milliseconds later. One Modbus IOModule, on a node with no route to the
coupler, cost 3.07s of the 15s to grant. That is the kernel's ARP clock, three
solicitations at a second each. A coupler that answers ARP and drops the
connection costs the transport's own timeout instead, 5s for Modbus and 10s
for OPC UA, and the loop paid it per IOModule in sequence. cmd/unit-runtime
opens the gRPC health server and the lease endpoint after Start returns, so a
unit with several modules on an unreachable segment would have spent tens of
seconds unable to be probed or granted. The same class as #1933 and #1935: a
network round trip on the goroutine the control plane depends on, bounded by
nothing the caller set.
The decision. A driver on a network protocol is dialled on a goroutine of
its own, and Start returns with the dial in flight. The outcome of the dial
was never a precondition of anything that follows it. A driver whose dial fails
already answers "will retry on first I/O", the scan's first read is that I/O,
and the outcome is logged from the goroutine in the same words it was logged in
before. The discriminator is the set wrapIfNetwork reads, because it is the
ReconnectingDriver the deferral leans on. The simulation driver is connected
before Start returns, as it was: its reads refuse until it is connected,
nothing re-dials it, and its connect touches no wire.
Two things in the wrapper had to change for the deferral to be safe, and the issue named the first of them as the thing to check before deciding.
The health monitor now starts whether or not the first dial succeeded. It used
to start on success only. So a device that was down when the runtime started
was dialled again by the next I/O and by nothing else, while a device that
dropped after a successful start was re-dialled every HealthInterval. Those
are one fault seen from two moments, and with the dial no longer waited on,
the first moment is the ordinary state of a replacement runtime on a node that
cannot yet reach its coupler. Both are watched the same way now.
The first dial runs under the wrapper's mutex, the way every reconnect always
has. A read that arrives while the first dial is on the wire waits for its
answer and finds the driver connected, instead of dialling beside it. And a
driver the adapter retires, on reload or on Stop, is closed: its
Disconnect waits for a dial in flight, and a dial that has not started yet,
from Connect or from a late reader's lazy reconnect, is refused with
ErrDriverClosed. Without that, a retirement landing before the deferred dial
started would have left a driver reconnecting a device nothing reads, with a
health monitor behind it that nothing would ever stop.
What the reload path does. reloadIOConfigs still dials a new driver
synchronously. It runs on the file watcher's goroutine, which nothing in the
control plane waits on, and a test reads the driver connected the instant the
reload returns. The issue's subject was Start, and the deferral stays there.
The tests. TestStart_ALeaseIsGrantedWhileTheCouplerHasNotAnswered gives
the runtime two IO drivers whose dials never answer and one simulation driver
that takes 300ms to connect, bounds Start at ten seconds, reads both dials in
flight and the simulation driver connected the moment Start returns, grants
the lease over the HTTP endpoint with both dials still on the wire, then lets
one coupler answer and reads that driver connected. Restoring the synchronous
dial fails it at the bound; deferring the simulation driver too fails it on the
300ms. In the driver package, TestReconnectingDriver_AFailedFirstDialIsStillWatched
refuses the first dial and reads the health monitor bring the driver up with
nothing else dialling. TestReconnectingDriver_AReadDuringTheFirstDialWaitsForIt
holds a dial on the wire, issues a read, and counts one dial and not two.
TestReconnectingDriver_ARetiredDriverDoesNotDial retires a driver before its
first dial and during it. Eight mutations, one per claim, each reddened the
test that carries it.
What is owed. On the next chart carrying this, the bench reading is the gap
between starting runtime and runtime started successfully with the coupler
unreachable, which should read in milliseconds, with the failed to connect IO
driver line arriving after gRPC health server listening.
Amendment (2026-09-22, #2151): a carried lease state names its epoch¶
The #1890 amendment taught the operator that an Expired is two findings
under one state name, and that the one where no renewal was ever acknowledged
at the binding is a replacement that failed to come up. Its test started from
an empty lease state, and the one road where a replacement exists never
reached that verdict. performFailover carries the predecessor's Expired
onto the binding it builds, so that the standby's first renewal is an
Expired to Held edge and writes the LeaseReestablished record #1308's
assembler closes the data gap on. noteLeaseState opened with
state == rb.LeaseState. The broker's own Expired for a standby that never
answered was equal to the carried one, so the operator wrote no log line, no
alarm and no audit record for it. The plant was told the primary had fenced
and that a failover was pending, and nothing about the runtime that was tried
and never took the lease.
The binding carries leaseStateEpoch beside leaseState now: the epoch the
broker observed the state at. A re-bind carries both values unchanged under
the new epoch, so a carried state is one whose epoch is behind the binding's,
and LeaseStateCarried in the API package is the one reading of that, shared
with the gateway. noteLeaseState treats a state equal to a carried one as a
transition, records it with #1890's start-up verdict, and stamps the epoch,
which is what keeps the next pass from recording it again. A state written
before the field existed has no epoch and is read as the binding's own, so an
upgrade does not re-record every unit that happened to be fenced at the time.
The audit reason stays LeaseExpired, as #1890 chose for the first-binding
case, and the batch-record assembler is taught the consequence: a gap opens
once. The second failover's opening expiry is the latest LeaseExpired at or
before its anchor walked back to the first of its run, where a run is broken
by a re-establishment or a runtime-restored record and bounded by the batch's
start. Without the walk the replacement's own record, which carries no
acknowledged contact because the replacement never answered, would have moved
the second event's opening off the instant the field went quiet and downgraded
it to Overlapping with a warning about a lease duration the samples stopped
long before.
The alarm is not doubled. The lease-expired alarm dedupes by prefix against
the active one, which still describes the primary's self-fence, and that
sentence is still true. The replacement's verdict reaches the log, the Part 11
record and the binding, and the gateway now reports Expired for that
binding because the broker declared it there, which #2150's reading of
leaseEstablishedBy alone could not tell from the carried case.
The tests. TestReconcileAvailability_AReplacementThatNeverComesUpIsRecorded
starts from the binding performFailover leaves and a runtime that refuses
every renewal, reads nothing recorded before the broker has spoken, one alarm
and one LeaseExpired record with the start-up verdict after it has, and
nothing more on a third pass. TestReconcileAvailability_ARecordedReplacementExpiryStillReestablishes
lets the standby answer late and reads the LeaseReestablished record behind
it. TestNoteLeaseState_AStateWithNoEpochIsThisBindingsOwn is the upgrade
guard. TestDataGapOpensOnceAcrossAReplacementThatNeverCameUp and
TestDataGapReopensAfterAReestablishment hold the assembler in both
directions. Restoring the equality check reddens the first and fourth;
removing the walk reddens the assembler test.