Skip to content

Edge-Runtime Failover Runbook

This runbook covers re-binding a unit's runtime to a designated standby node when its edge controller is lost: the operational procedure for the hold-then-resume failover mechanism decided in ADR 0006 and implemented across the API, runtime, operator, CLI/gateway, and UI.

Failover moves the software runtime (the FB scan loop and I/O driver connections) to other hardware. The controller↔I/O link is a network connection (remote Ethernet I/O), so the standby needs field-network reach but no physical adjacency. It does not replicate live state: recovery is an ISA-88 Restart, and field outputs hold last value (or device fail-safe) during the gap. A runtime that restarts in place restores the commanded values and block operating points it was carrying before its first scan. The gap ends on a continuation, where it used to end on a compile-time default (ADR 0080).

A promoted standby remembers nothing about this plant, so it resumes from the plant (ADR 0082). Each device output is read back at the address its IOModule channel declares in readbackAddress, and the block resumes on the value the coupler is holding. Where no channel declares one, that output writes nothing until the Restart's own restarting logic commands it, and the coupler goes on holding. Declaring readbackAddress is what buys the continuation. Discovery writes it automatically. A hand-authored IOModule does not have it unless somebody put it there. dcs_runtime_output_adoption_total reports no-readback once per channel per promotion, so a promql check after a drill says which side of this a deployment is on.

Before you start

  • Know the unit's availability mode. dcs get unit <name> -s <site> shows it, or the System app → Controller detail → Hosted Units table (Availability column). The procedure differs by mode (below).
  • Confirm what the unit has armed at the edge. dcs get units -s <site> carries an EDGE HOLD column, and the Unit detail page carries the same verdict as its Edge Hold row (#1833). A Failover unit that reads nothing armed runs no safe-state chart when its lease expires, and one that reads armed, cannot fire has already spent its Autonomy watchdog in this runtime process. Both are worth knowing before you take the node away, and neither is visible in any other cell on the row.
  • Confirm a standby exists. dcs unit failover <name> -s <site> (no --to-node) lists the eligible targets the operator resolved from spec.availability.failoverTargets. An empty list means no enrolled, Ready, field-reachable standby is configured. Fix that first (see production-deployment.md).
  • A manual failover requires the unit not be actively Running. A dead-node unit is normally already Held, whether by the watchdog/self-hold path or (Failover mode) by the automatic path itself once the lease expiry outlasts the hold bound. If it is still Running, Hold the owning batch first (dcs command Batch <batch> Hold -s <site> --reason "…").
  • A failover during a batch lands in that batch's production record. The record carries the control gap and the data gap separately, each bound naming the AuditRecord that evidences it (Batch Production Records). A manual failover opens the data gap at the last heartbeat or lease renewal the control plane had acknowledged from the old runtime. That is earlier than the re-bind, and it is the last instant the product itself observed that runtime alive (#1753, #1805). It is still not the instant you fenced the node. Put the fencing time in --reason, where a reviewer will find it. Where the product never reached that runtime at all the bound falls back to the re-bind. The record discloses that on its PartialData condition and the batch lands at PendingReview.

Mode Autonomy (default) — manual fenced failover

In Autonomy mode the edge keeps controlling through a control-plane partition, so the operator must physically fence the old node before re-binding. Otherwise two writers could drive the same field devices.

  1. Physically fence the old node. Power it off, or disconnect it from the field network. This is the load-bearing step: confirm it, do not assume it. On Modbus there is no session exclusivity. A live old node and a new one both writing is undetected, last-write-wins.

    On a bonded controller, the field network is two cables

    The recommended posture bonds the field interface across two ports on two switches (ADR 0026). Pulling one member fails the bond over to the other and the node keeps writing. A half-fenced node is exactly the dual-writer hazard this step exists to prevent. Pull both members, or power the node off. Powering off is the unambiguous fence. Prefer it.

    1. List eligible targets: dcs unit failover <name> -s <site>.
    2. Re-bind, certifying the fence:
      dcs unit failover <name> -s <site> --to-node <standby> \
        --confirm-fenced --reason "primary IPC PSU failure, node powered off"
      
      --confirm-fenced is required in Autonomy mode and is recorded on the AuditRecord (21 CFR Part 11). Omitting it is rejected. --reason is required on every failover path (#687) and must meet the deployment reason policy.
    3. If the standby already hosts another unit's runtime, the co-location guard rejects the request. Add --allow-colocation only after accepting the 1:1-controller:unit trade-off.

UI equivalent: System app → Controller detail → Hosted Units → Failover (visible to operate-lead). The dialog's fenced-confirmation checkbox and justification are required and the target picker shows co-location occupancy.

Steps 2 to 4 on a two-node edge pair: the standby listed as occupied before anything is submitted, the unfenced re-bind rejected, the co-located re-bind rejected, and then the same request with the fence certified and the trade-off accepted, landing with a new fencing epoch. The node is fenced by the operator off camera. Nothing here failed on its own.

Mode Failover (opt-in) — lease-based

In Failover mode the runtime holds an operator-brokered control lease and self-fences when it cannot renew (FB output writes stop, reads continue). The decision runs on its own local clock, so it works under a control-plane partition. The operator then re-binds automatically.

What the runtime dates that clock from is worth knowing when you are reading timings off a partitioned pair. Every renewal carries the operator's own acknowledged-renewal age, and the runtime expires against that age. Both sides are therefore counting from the same moment: the last renewal the operator had acknowledged. A partition that delivers the operator's POSTs while dropping the runtime's replies buys the runtime no extra time, and the runtime has begun its bounded hold by the time the operator declares the lease Expired (#1802).

Automatic path (normal case): no operator action. Once the operator has observed the lease Expired for the safety margin (½ the lease duration, clamped 10–60s), it re-binds the runtime to the first free eligible standby, stamps a new epoch, and recreates the pod there. If the unit is still Running when the lease expires (a batch mid-run when the node died), the operator first places it on ISA-88 Hold itself once the expiry outlasts the hold bound (½ the lease, clamped 5–30s). The result is the same Hold, Critical alarm, and RuntimeCrashDetected condition the runtime crash detector produces. After the margin it re-binds (hold-then-resume, ADR 0006). A runtime that recovers and renews its lease inside the hold bound rides through with no Hold. Watch it land:

dcs get unit <name> -s <site>   # RuntimeBinding + FailoverRequest conditions

The UI answers the same question without a reload. The Unit detail page and the Controller detail's Hosted Units table both re-read every ten seconds. The state badge, the Availability row and the lease column therefore move on their own while you watch. A unit re-bound to its standby leaves the old chassis's Hosted Units table and appears in the new one.

The automatic path with nothing for an operator to do but watch. The control lease expires and the runtime self-fences, the operator re-binds it to the standby, and only afterwards does Kubernetes mark the lost machine Offline. The kill is injected off camera. What is filmed is the lease and the two chassis on either side of it.

When the automatic path is blocked: if every eligible target is occupied, a NoFreeTarget condition is set and no re-bind happens. Either free/provision a standby, or run a manual co-located failover:

dcs unit failover <name> -s <site> --to-node <standby> --allow-colocation \
  --reason "no free standby; co-locating during incident <ref>"

When the automatic path suspends itself: the automatic path may re-bind a unit to each free eligible standby once between confirmed runtimes. When every one of them has been tried and none took the control lease, it stops, sets a FailoverSuspended condition and raises a Critical alarm naming the standbys it tried (#1890). dcs get unit <name> -s <site> reports it as autoFailover.suspended with the standbys under autoFailover.triedNodes, and the Unit detail page marks the Availability row. The nodes are on the object itself as status.runtimeBinding.unsettledRebinds.

Nothing needs to be done to resume. The pod on the current binding is still there, and the operator is still renewing against it. The first renewal that runtime acknowledges clears the condition, the alarm and the ledger. What the suspension is telling you is that re-binding is not the remedy: no replacement is coming up anywhere, which is a registry, an image, a driver endpoint, or a start-up slower than availability.runtimeStartupGraceSeconds. A message broker the replacement cannot reach was on that list until #1933. A runtime now starts, opens its lease endpoint and is granted its lease with the broker still away. A coupler the replacement cannot reach was on it until #1935. The promoted outputs now ask the plant beside the scan, and the lease endpoint is up before the coupler has answered. The drivers' own first dial was the last of the three, until #1936. A network driver now dials on its own goroutine, and a dial that failed is retried by the first I/O and by the driver's health monitor. Read kubectl describe pod <unit>-runtime -n site-<site> before doing anything to the failover configuration.

A deliberate failover is available throughout and empties the budget, so an operator who has a target in mind is never blocked by the suspension.

When the automatic path is waiting on a recent leadership change: a FailoverRequest condition with reason LeadershipHoldOff means the re-bind is delayed and nothing is wrong. A physical-operator that loses the Kubernetes API does not exit any more. It stands down for physicalOperator.leaseStandDown (90s by default) and goes on renewing the leases it already holds, because a grant of liveness travels to the runtime's pod IP and never touches the apiserver (#1894). The operator that takes leadership next therefore cannot treat a lease expiry as proof that the old runtime has stopped writing until that window, the unit's lease duration and the re-bind margin have all elapsed since it acquired leadership. The condition message names the instant the wait ends, and the re-bind proceeds on its own when the wait is over. If the lease comes back Held instead, because the predecessor really was still granting it, nothing re-binds and the operator removes the condition on the first pass after the deadline (#2148). A LeadershipHoldOff you can still read is therefore a wait that is still running.

Nothing needs to be done, and there is one thing not to do: do not shorten physicalOperator.leaseStandDown to make the wait go away. It is what keeps a platform event on the control-plane side from fencing units whose nodes were never touched. A failover that genuinely cannot wait is a manual one with --confirm-fenced. That road is not held off at all. Powering a node off is a stronger fence than any arithmetic here can construct.

When a control-plane node has just lost power and a unit fenced anyway: the standby physical-operator may have been pinned to that node's apiserver. Its connection is not closed by a node losing power, and client-go keeps using it until its HTTP/2 health check gives up. Since #1930 that check is 2s and 2s on this operator (physicalOperator.apiserverHealthCheck), so the standby dials a live apiserver inside 4s and the handoff fits the runtime's budget. A fence on that fault is a defect to report with the standby's log, from its last Error retrieving lease lock to Successfully acquired lease, and not a wait to sit out.

Which expiry a lease-expired alarm is reporting: the alarm says so in its own sentence. self-fenced is a runtime that held the lease and stopped answering, and the standby takes over from a writer that has stood down. never established is a runtime that was bound and never took the lease, so nothing was driving that unit and nobody fenced. The second one points at the start-up. The alarm names the grace the runtime was given along with what its last renewal returned. A refusal came from the runtime's own handler, which means that runtime is up and answering. A transport error means it was never reached.

The verdict survives a control-plane leader change, and it is worth knowing why. The operator learns which of the two it is from whether a renewal has ever been acknowledged at that binding, and a replacement physical-operator did not observe the ones its predecessor got. It reads status.runtimeBinding.leaseEstablishedBy instead, which names the runtime that answered. So an expiry a few minutes after an operator restart is read against the runtime's own history and not against the new process's. A never established alarm still means what it says. That runtime has answered nobody, under any operator.

Planned maintenance (lease still Held): to move a healthy Failover-mode unit (e.g. to service its node), the lease has not expired, so use the fenced-confirmation override exactly as in Autonomy mode. Fence the node first, then --confirm-fenced.

After failover (both modes)

  1. Verify the new binding: dcs get unit <name> -s <site> shows RuntimeBinding.node = the standby and an incremented epoch. For Failover mode the lease returns to Held once renewals resume.
  2. The control program redeploys automatically from the control-plane source of truth (CRDs). There is no hostPath migration. Confirm the runtime pod is Ready and drivers reconnected (Diagnostics page, or dcs get unit <name> -s <site>).
  3. Confirm the runtime is driving the field, and not merely permitted to. A Ready pod is permitted to write, and in Failover mode a lease back at Held says so too. Neither is evidence that anything reached the plant. A promoted output resumes on the value its coupler is holding, and one whose channel declares no readbackAddress withholds its writes until something commands it (ADR 0082). Such a runtime is healthy, permitted, and driving nothing.

dcs_runtime_outputs_unestablished says how many outputs are still withholding, in either mode. A count above zero is answered by step 4, because the recipe's restarting logic is what commands them.

Read dcs_runtime_outputs_holding beside it, and do not stop at the first gauge. An output whose channel does declare a readbackAddress resumes on the value its coupler is holding and then writes that value every scan, ignoring its own input, until something commands it. That is the good path and it is still not control: the loop behind that address is not regulating, and every counter below reads exactly as it does on a runtime that is (#1816). The unestablished gauge counts the other phase and reads 0. Step 4 answers this count too, and this is the count that stays up if the recipe's restarting logic never writes the setpoint.

In Failover mode there is a second signal, and it is the sharper one. The runtime logs one line per fence episode, at the first output write that reached the device, naming the address and how long it sat between the unfence and that write:

runtime resumed driving the field: first output write since the fence lifted

dcs_runtime_writes_resumed_total carries the same edge as a counter. An Autonomy runtime has no fence and emits neither, so read the gauge there.

dcs_runtime_writes_total says the same thing continuously rather than once. Its fenced outcome counts the writes the fence dropped and its ok outcome counts only the writes that reached a device, so a promoted runtime that is driving shows a flat fenced rate and a moving ok rate. Both rates flat is the withholding case above, and it is the reason to read the gauge next. A moving ok rate is not evidence of control, because a holding output writes: read dcs_runtime_outputs_holding, and dcs get runtime -u <name> for the block and address behind a non-zero count. 4. Restart the process. Recovery is an ISA-88 Restart. Issue it through the batch (dcs command Batch <batch> Restart -s <site>) and confirm the recipe-defined restarting logic re-establishes process conditions. 5. Clear the alarms. Both Alarms follow the runtime, and neither waits for step 4. The lease-expired (critical) Alarm clears when the lease re-establishes on the standby. The failover (high) Alarm says the runtime was re-bound and is not back to normal yet, so it clears when the replacement pod is Ready and holding its control lease on the new node (#1733). Both stamp status.clearedAt. Both also keep whatever acknowledgement they already carried, so a signature you have already given is not asked for twice. Acknowledge per your ISA-18.2 workflow.

A failover Alarm still sitting Active therefore means the runtime has not come back. That is a live problem, and it is worth reading before step 4.

The two intervals overlap, and it is worth knowing which part each covers. The lease-expired Alarm opens at lease expiry. Its span covers detection, the safety margin, the re-bind and the lease coming back. The failover Alarm opens at the re-bind and closes when the replacement pod goes Ready. In Failover mode the runtime starts fenced, and it stays fenced until the first lease renewal reaches it. A fenced runtime fails its readiness probe. The lease grant is therefore already behind you by the time this Alarm clears. Read the span as the time to a Ready pod on the standby, and do not apportion it between pod start and the lease. Neither one measures how long the batch was held, because the ISA-88 Restart is the operator's and no alarm waits for it.

One set of failover Alarms does not carry that reading. Versions before 0.5.1 raised them and cleared none, so a cluster upgrading from one arrives holding whatever it accumulated. The upgrade does not clear them, because the clear is driven by a failover and those are already over. The next failover on that unit sweeps them alongside its own, and their status.clearedAt is that sweep. On an inherited Alarm the interval is therefore the age of the annunciation and not a recovery time. Check the raise timestamps against your upgrade before reading any of them as a duration.

Failing back

There is no automatic fail-back. The binding is status-owned (RuntimeBinding.directedNode) and survives a GitOps sync of the spec. Once the original node is repaired and re-enrolled, fail back with another explicit failover to it (same procedure, naming the original node).