Edge-Runtime Failover Runbook¶
This runbook covers re-binding a unit's runtime to a designated standby node when its edge controller is lost: the operational procedure for the hold-then-resume failover mechanism decided in ADR 0006 and implemented across the API, runtime, operator, CLI/gateway, and UI.
Failover moves the software runtime (the FB scan loop and I/O driver connections) to other hardware. The controller↔I/O link is a network connection (remote Ethernet I/O), so the standby needs field-network reach but no physical adjacency. It does not replicate live state: recovery is an ISA-88 Restart, and field outputs hold last value (or device fail-safe) during the gap. A runtime that restarts in place restores the commanded values and block operating points it was carrying before its first scan. The gap ends on a continuation, where it used to end on a compile-time default (ADR 0080).
A promoted standby remembers nothing about this plant, so it resumes from the
plant
(ADR 0082). Each
device output is read back at the address its IOModule channel declares in
readbackAddress, and the block resumes on the value the coupler is holding.
Where no channel declares one, that output writes nothing until the Restart's
own restarting logic commands it, and the coupler goes on holding. Declaring
readbackAddress is what buys the continuation. Discovery writes it
automatically. A hand-authored IOModule does not have it unless somebody put it
there. dcs_runtime_output_adoption_total reports no-readback once per
channel per promotion, so a promql check after a drill says which side of
this a deployment is on.
Before you start¶
- Know the unit's availability mode.
dcs get unit <name> -s <site>shows it, or the System app → Controller detail → Hosted Units table (Availability column). The procedure differs by mode (below). - Confirm what the unit has armed at the edge.
dcs get units -s <site>carries an EDGE HOLD column, and the Unit detail page carries the same verdict as its Edge Hold row (#1833). AFailoverunit that readsnothing armedruns no safe-state chart when its lease expires, and one that readsarmed, cannot firehas already spent itsAutonomywatchdog in this runtime process. Both are worth knowing before you take the node away, and neither is visible in any other cell on the row. - Confirm a standby exists.
dcs unit failover <name> -s <site>(no--to-node) lists the eligible targets the operator resolved fromspec.availability.failoverTargets. An empty list means no enrolled, Ready, field-reachable standby is configured. Fix that first (see production-deployment.md). - A manual failover requires the unit not be actively Running. A
dead-node unit is normally already Held, whether by the watchdog/self-hold
path or (
Failovermode) by the automatic path itself once the lease expiry outlasts the hold bound. If it is still Running, Hold the owning batch first (dcs command Batch <batch> Hold -s <site> --reason "…"). - A failover during a batch lands in that batch's production record.
The record carries the control gap and the data gap separately, each
bound naming the AuditRecord that evidences it
(Batch Production Records).
A manual failover opens the data gap at the last heartbeat or lease
renewal the control plane had acknowledged from the old runtime. That is
earlier than the re-bind, and it is the last instant the product itself
observed that runtime alive (#1753, #1805). It is still not the instant
you fenced the node. Put the fencing time in
--reason, where a reviewer will find it. Where the product never reached that runtime at all the bound falls back to the re-bind. The record discloses that on itsPartialDatacondition and the batch lands atPendingReview.
Mode Autonomy (default) — manual fenced failover¶
In Autonomy mode the edge keeps controlling through a control-plane
partition, so the operator must physically fence the old node before
re-binding. Otherwise two writers could drive the same field devices.
-
Physically fence the old node. Power it off, or disconnect it from the field network. This is the load-bearing step: confirm it, do not assume it. On Modbus there is no session exclusivity. A live old node and a new one both writing is undetected, last-write-wins.
On a bonded controller, the field network is two cables
The recommended posture bonds the field interface across two ports on two switches (ADR 0026). Pulling one member fails the bond over to the other and the node keeps writing. A half-fenced node is exactly the dual-writer hazard this step exists to prevent. Pull both members, or power the node off. Powering off is the unambiguous fence. Prefer it.
- List eligible targets:
dcs unit failover <name> -s <site>. - Re-bind, certifying the fence:
dcs unit failover <name> -s <site> --to-node <standby> \ --confirm-fenced --reason "primary IPC PSU failure, node powered off"--confirm-fencedis required inAutonomymode and is recorded on the AuditRecord (21 CFR Part 11). Omitting it is rejected.--reasonis required on every failover path (#687) and must meet the deployment reason policy. - If the standby already hosts another unit's runtime, the co-location
guard rejects the request. Add
--allow-colocationonly after accepting the 1:1-controller:unit trade-off.
- List eligible targets:
UI equivalent: System app → Controller detail → Hosted Units → Failover
(visible to operate-lead). The dialog's fenced-confirmation checkbox and
justification are required and the target picker shows co-location
occupancy.
Mode Failover (opt-in) — lease-based¶
In Failover mode the runtime holds an operator-brokered control lease
and self-fences when it cannot renew (FB output writes stop, reads
continue). The decision runs on its own local clock, so it works under a
control-plane partition. The operator then re-binds automatically.
What the runtime dates that clock from is worth knowing when you are
reading timings off a partitioned pair. Every renewal carries the
operator's own acknowledged-renewal age, and the runtime expires against
that age. Both sides are therefore counting from the same moment: the
last renewal the operator had acknowledged. A partition that delivers the
operator's POSTs while dropping the runtime's replies buys the runtime no
extra time, and the runtime has begun its bounded hold by the time the
operator declares the lease Expired (#1802).
Automatic path (normal case): no operator action. Once the operator
has observed the lease Expired for the safety margin (½ the lease
duration, clamped 10–60s), it re-binds the runtime to the first free
eligible standby, stamps a new epoch, and recreates the pod there. If the
unit is still Running when the lease expires (a batch mid-run when the
node died), the operator first places it on ISA-88 Hold itself once the
expiry outlasts the hold bound (½ the lease, clamped 5–30s). The result
is the same Hold, Critical alarm, and RuntimeCrashDetected condition
the runtime crash detector produces. After the margin it re-binds
(hold-then-resume, ADR 0006). A runtime that recovers and renews its
lease inside the hold bound rides through with no Hold. Watch it land:
dcs get unit <name> -s <site> # RuntimeBinding + FailoverRequest conditions
The UI answers the same question without a reload. The Unit detail page and the Controller detail's Hosted Units table both re-read every ten seconds. The state badge, the Availability row and the lease column therefore move on their own while you watch. A unit re-bound to its standby leaves the old chassis's Hosted Units table and appears in the new one.
When the automatic path is blocked: if every eligible target is
occupied, a NoFreeTarget condition is set and no re-bind happens. Either
free/provision a standby, or run a manual co-located failover:
dcs unit failover <name> -s <site> --to-node <standby> --allow-colocation \
--reason "no free standby; co-locating during incident <ref>"
When the automatic path suspends itself: the automatic path may re-bind a
unit to each free eligible standby once between confirmed runtimes. When every
one of them has been tried and none took the control lease, it stops, sets a
FailoverSuspended condition and raises a Critical alarm naming the standbys
it tried
(#1890).
dcs get unit <name> -s <site> reports it as autoFailover.suspended with the
standbys under autoFailover.triedNodes, and the Unit detail page marks the
Availability row. The nodes are on the object itself as
status.runtimeBinding.unsettledRebinds.
Nothing needs to be done to resume. The pod on the current binding is still
there, and the operator is still renewing against it. The first renewal that
runtime acknowledges clears the condition, the alarm and the ledger. What the
suspension is telling you is that re-binding is not the remedy: no replacement
is coming up anywhere, which is a registry, an image, a driver endpoint, or a
start-up slower than availability.runtimeStartupGraceSeconds. A message
broker the replacement cannot reach was on that list until
#1933. A
runtime now starts, opens its lease endpoint and is granted its lease with the
broker still away. A coupler the replacement cannot reach was on it until
#1935. The
promoted outputs now ask the plant beside the scan, and the lease endpoint is
up before the coupler has answered. The drivers' own first dial was the last
of the three, until
#1936. A
network driver now dials on its own goroutine, and a dial that failed is
retried by the first I/O and by the driver's health monitor. Read
kubectl describe pod <unit>-runtime -n site-<site> before doing anything to
the failover configuration.
A deliberate failover is available throughout and empties the budget, so an operator who has a target in mind is never blocked by the suspension.
When the automatic path is waiting on a recent leadership change: a
FailoverRequest condition with reason LeadershipHoldOff means the re-bind is
delayed and nothing is wrong. A physical-operator that loses the Kubernetes
API does not exit any more. It stands down for physicalOperator.leaseStandDown
(90s by default) and goes on renewing the leases it already holds, because a
grant of liveness travels to the runtime's pod IP and never touches the
apiserver
(#1894).
The operator that takes leadership next therefore cannot treat a lease expiry as
proof that the old runtime has stopped writing until that window, the unit's
lease duration and the re-bind margin have all elapsed since it acquired
leadership. The condition message names the instant the wait ends, and the
re-bind proceeds on its own when the wait is over. If the lease comes back
Held instead, because the predecessor really was still granting it, nothing
re-binds and the operator removes the condition on the first pass after the
deadline
(#2148).
A LeadershipHoldOff you can still read is therefore a wait that is still
running.
Nothing needs to be done, and there is one thing not to do: do not shorten
physicalOperator.leaseStandDown to make the wait go away. It is what keeps a
platform event on the control-plane side from fencing units whose nodes were
never touched. A failover that genuinely cannot wait is a manual one with
--confirm-fenced. That road is not held off at all. Powering a node off is a
stronger fence than any arithmetic here can construct.
When a control-plane node has just lost power and a unit fenced anyway:
the standby physical-operator may have been pinned to that node's apiserver.
Its connection is not closed by a node losing power, and client-go keeps using
it until its HTTP/2 health check gives up. Since
#1930 that
check is 2s and 2s on this operator (physicalOperator.apiserverHealthCheck),
so the standby dials a live apiserver inside 4s and the handoff fits the
runtime's budget. A fence on that fault is a defect to report with the
standby's log, from its last Error retrieving lease lock to
Successfully acquired lease, and not a wait to sit out.
Which expiry a lease-expired alarm is reporting: the alarm says so in its
own sentence. self-fenced is a runtime that held the lease and stopped
answering, and the standby takes over from a writer that has stood down.
never established is a runtime that was bound and never took the lease, so
nothing was driving that unit and nobody fenced. The second one points at the
start-up. The alarm names the grace the runtime was given
along with what its last renewal returned. A refusal came from the runtime's
own handler, which means that runtime is up and answering. A transport error
means it was never reached.
The verdict survives a control-plane leader change, and it is worth knowing
why. The operator learns which of the two it is from whether a renewal has ever
been acknowledged at that binding, and a replacement physical-operator did
not observe the ones its predecessor got. It reads
status.runtimeBinding.leaseEstablishedBy instead, which names the runtime
that answered. So an expiry a few minutes after an operator restart is read
against the runtime's own history and not against the new process's. A
never established alarm still means what it says. That runtime has answered
nobody, under any operator.
Planned maintenance (lease still Held): to move a healthy
Failover-mode unit (e.g. to service its node), the lease has not
expired, so use the fenced-confirmation override exactly as in Autonomy
mode. Fence the node first, then --confirm-fenced.
After failover (both modes)¶
- Verify the new binding:
dcs get unit <name> -s <site>showsRuntimeBinding.node= the standby and an incremented epoch. ForFailovermode the lease returns toHeldonce renewals resume. - The control program redeploys automatically from the control-plane
source of truth (CRDs). There is no hostPath migration. Confirm the
runtime pod is Ready and drivers reconnected (Diagnostics page, or
dcs get unit <name> -s <site>). - Confirm the runtime is driving the field, and not merely permitted
to. A Ready pod is permitted to write, and in
Failovermode a lease back atHeldsays so too. Neither is evidence that anything reached the plant. A promoted output resumes on the value its coupler is holding, and one whose channel declares noreadbackAddresswithholds its writes until something commands it (ADR 0082). Such a runtime is healthy, permitted, and driving nothing.
dcs_runtime_outputs_unestablished says how many outputs are still
withholding, in either mode. A count above zero is answered by step 4,
because the recipe's restarting logic is what commands them.
Read dcs_runtime_outputs_holding beside it, and do not stop at the
first gauge. An output whose channel does declare a readbackAddress
resumes on the value its coupler is holding and then writes that value
every scan, ignoring its own input, until something commands it. That is
the good path and it is still not control: the loop behind that address is
not regulating, and every counter below reads exactly as it does on a
runtime that is
(#1816).
The unestablished gauge counts the other phase and reads 0. Step 4 answers
this count too, and this is the count that stays up if the recipe's
restarting logic never writes the setpoint.
In Failover mode there is a second signal, and it is the sharper one. The
runtime logs one line per fence episode, at the first output write that
reached the device, naming the address and how long it sat between the
unfence and that write:
runtime resumed driving the field: first output write since the fence lifted
dcs_runtime_writes_resumed_total carries the same edge as a counter. An
Autonomy runtime has no fence and emits neither, so read the gauge there.
dcs_runtime_writes_total says the same thing continuously rather than
once. Its fenced outcome counts the writes the fence dropped and its ok
outcome counts only the writes that reached a device, so a promoted runtime
that is driving shows a flat fenced rate and a moving ok rate. Both
rates flat is the withholding case above, and it is the reason to read the
gauge next. A moving ok rate is not evidence of control, because a
holding output writes: read dcs_runtime_outputs_holding, and
dcs get runtime -u <name> for the block and address behind a non-zero
count.
4. Restart the process. Recovery is an ISA-88 Restart. Issue it
through the batch (dcs command Batch <batch> Restart -s <site>) and
confirm the recipe-defined restarting logic re-establishes process
conditions.
5. Clear the alarms. Both Alarms follow the runtime, and neither waits for
step 4. The lease-expired (critical) Alarm clears when the lease
re-establishes on the standby. The failover (high) Alarm says the runtime
was re-bound and is not back to normal yet, so it clears when the
replacement pod is Ready and holding its control lease on the new node
(#1733). Both stamp status.clearedAt. Both also keep whatever
acknowledgement they already carried, so a signature you have already given
is not asked for twice. Acknowledge per your ISA-18.2 workflow.
A failover Alarm still sitting Active therefore means the runtime has not come back. That is a live problem, and it is worth reading before step 4.
The two intervals overlap, and it is worth knowing which part each covers.
The lease-expired Alarm opens at lease expiry. Its span covers detection,
the safety margin, the re-bind and the lease coming back. The failover Alarm
opens at the re-bind and closes when the replacement pod goes Ready. In
Failover mode the runtime starts fenced, and it stays fenced until the
first lease renewal reaches it. A fenced runtime fails its readiness probe.
The lease grant is therefore already behind you by the time this Alarm
clears. Read the span as the time to a Ready pod on the standby, and do not
apportion it between pod start and the lease. Neither one measures how long
the batch was held, because the ISA-88 Restart is the operator's and no
alarm waits for it.
One set of failover Alarms does not carry that reading. Versions before
0.5.1 raised them and cleared none, so a cluster upgrading from one arrives
holding whatever it accumulated. The upgrade does not clear them, because
the clear is driven by a failover and those are already over. The next
failover on that unit sweeps them alongside its own, and their
status.clearedAt is that sweep. On an inherited Alarm the interval is
therefore the age of the annunciation and not a recovery time. Check the
raise timestamps against your upgrade before reading any of them as a
duration.
Failing back¶
There is no automatic fail-back. The binding is status-owned
(RuntimeBinding.directedNode) and survives a GitOps sync of the spec.
Once the original node is repaired and re-enrolled, fail back with another
explicit failover to it (same procedure, naming the original node).
Related Documentation¶
- High Availability and Failure Modes — the fenced/lease alert runbook and failure-mode table
- Reference Architectures — Pattern B standby topology (dedicated vs shared)
- Production Deployment — provisioning and labeling standby edge nodes
- ADR 0006 — the decision and its rationale