ADR 0053: An IOModule the probe could not measure is Unknown, not Offline¶
Status: Accepted Date: 2026-08-11 Issue: #1477
Amended (2026-08-21, #1719): the retained
Offlinecascade splits by protocol. Losing the Controller is a determined loss of the device only forprotocol: simulation, whose liveness is the Controller's heartbeat. For a network protocol the Controller is only where the probe runs, so the same cascade now readsUnknown. See Amendment at the end of this document.
Context¶
IOModule.status.state is derived from an io-probe pod: a per-Controller one for
a module that names a controllerRef, or a namespace-shared one for a
controller-less network module (ADR 0021,
ADR 0042). The probe
answers two things about a device, and the reconciler mapped them faithfully:
connected became Online, and reachable-but-not-connected became Fault.
Every other outcome also produced a state, and that is where the mapping stopped
being faithful. A probe pod that was missing, that had no IP, that did not answer
the query, or that had not yet loaded the module into its config all produced
Offline — a claim about the device, made at the moment the DCS had lost the
only instrument it had for making claims about the device.
The bench found the cost on 2026-08-11. A node shutdown left the probe pod in a
terminal phase, nothing replaced it, and wago-750-352 read Offline for
twenty-three hours. The coupler was answering the whole time: the unit runtime
held its own Modbus connection and published input.0 and input.1 every
second. An operator reading Offline, or reading the dcs io read error that
named the probe pod's missing IP, is pointed at the network and the hardware.
Both were fine.
Offline is not only a display value. Fault and Offline are the auto-HOLD
set: on the transition into either from a non-degraded state, the IOModule
reconciler annotates every Running unit allocated to a batch with Hold
(ISA-88 Clause 7.4, equipment fault → automatic procedural response). So a
monitoring sidecar dying could hold production on equipment that was running
correctly, and the io-probe is exactly that — a sidecar. It commissions and
watches; it is not in the control path. The runtime that drives the unit reaches
the device over its own connection and goes on scanning while the probe is down.
Unknown already existed in the enum and already carried this meaning
everywhere except in the code. docs/io.md documents a newly created module
transitioning Unknown → Online, and the gateway maps an empty state to
Unknown when it renders. Only the reconciler never produced it.
Decision¶
Offline is a measurement. Unknown is the absence of one.
The reconciler returns Unknown for every branch in which the probe did not
answer for the module: no probe pod name, a pod that is missing, a pod in a
terminal phase, a pod with no IP, a query that errored, and a module the probe
has not loaded yet.
Offline is kept for the cascade that is genuinely determined — the Controller
is missing, is not Joined, its node is not Ready, or a simulation
Controller's heartbeat has gone stale. In each of those the DCS has no path to
the device and knows it.
Fault is unchanged. It is the probe reaching the device's endpoint and being
refused, which is a measurement and the one the auto-HOLD exists for.
Two consequences follow directly and are part of the decision:
Unknownis outside the auto-HOLD set. Losing the instrument is not an equipment fault, and holding a batch on it is a false trip.Unknownis annunciated. The alarm generator raises the sameAlarmTypeSystemalarm at the sameMediumseverity that theOfflinebranch raised, with a message naming the probe rather than a communication loss. Without this the correction would have traded a wrong annunciation for no annunciation, which is worse: nothing would say the plant had lost sight of a device.
Alternatives Considered¶
- Leave the mapping and fix only the pod replacement. The pod-replacement
bug in #1477 is real and is fixed independently: a probe pod in a terminal
phase is now treated as absent and recreated. But the mapping is wrong on its
own terms and has other ways to be reached — a probe still starting, a probe
unschedulable on a zoned OT network (ADR 0042),
a config that has not propagated. Every one of them would still have reported
a healthy device as
Offlineand could still have held a batch. - Derive state from the unit runtime when one is scanning the device. The
runtime is the strongest available evidence that a device is answering, and it
is what made the bench diagnosis so slow — the DCS held the refuting evidence
the whole time. Rejected here as a larger change than the correction needs: it
couples IOModule state to Unit lifecycle, which
ADR 0021 rejected for the
connectivity signal, and a module may be referenced by ControlModules under
several Units or under none.
Unknownstates the position honestly without inventing a second source of truth. Using the runtime as corroboration remains open. - A new state for "probe unavailable". Rejected as a needless enum widening.
Unknownmeans exactly this and is already the documented pre-measurement state. - Keep
Unknownunannunciated, on the grounds that it is the normal startup state. Rejected. A module sits inUnknownfor the seconds its probe takes to start, which is the same window that raised anOfflinealarm before this change, so the noise is unchanged. Staying silent afterwards would make a plant-wide loss of I/O visibility the quietest thing on the alarm list.
Note that ADR 0021 rejected an
alternative described as "skip the probe for controller-less IOModules (state
Unknown)". That is a different proposal. It removed the probe, so the module
could never reach Fault and a device dropping mid-batch would never HOLD. Here
the probe is unchanged and Fault still HOLDs. What changes is only what the DCS
says when the probe itself cannot answer.
Consequences¶
- An
AlarmDefinitionwritten againstStateEquals: Offlineno longer fires for a dead or starting probe. That is the intent, and a deployment that wants to act on lost visibility adds aStateEquals: Unknowndefinition. Definitions that carryexceptionAction: Holdshould stay onFaultandOffline, where the DCS has actually observed something about the equipment. kubectl get iomodulesshowsUnknownin cases that used to printOffline. The state now answers a different question — what the DCS knows, not what it guesses — so a runbook that readsOfflineas "check the network" is correct for the first time.- The
Offlinebranch of the auto-HOLD narrows to the cases where the Controller or its node is gone. The ISA-88 Clause 7.4 traceability indocs/compliance/isa88.mdis unchanged in substance: a device that drops mid-batch still readsFaultfrom a working probe and still holds the Unit. Unknownbecoming reachable means the state can flapUnknown → Onlineduring a probe restart. That is the honest sequence, and it is the onedocs/io.mdalready documented for a newly created module.
Amendment (2026-08-21, #1719): the retained cascade splits by protocol¶
This ADR moved every probe-side branch to Unknown and kept Offline for
"the cascade that is genuinely determined — the Controller is missing, is not
Joined, its node is not Ready, or a simulation Controller's heartbeat has
gone stale." That sentence bundles two different things under one word.
For protocol: simulation it holds exactly. The module has no endpoint, its
Online state is derived from the Controller's Status.LastAlive, and the
Controller is the only instrument it has. Losing it is a determined loss of the
device.
For a network protocol it does not hold. controllerRef is inert on the
tag-read path (ADR 0021): the
runtime resolves each tag from its ioModuleRef and address, opens its own
connection, and goes on scanning while the named chassis is dark. The
Controller's only role for such a module is carrying the io-probe pod. Reading
Offline there asserts that the device is gone on the strength of having lost
the host of a monitoring sidecar, which is the claim this ADR was written to
refuse.
The bench walked into it on 2026-08-20 during the #942
HA drills. wago-750-352 went Offline the moment its controllerRef node was
powered off, while the coupler answered the unit runtime on the standby chassis
throughout. That is the same reading, and the same wrong diagnosis, that this
ADR's own Context records from 2026-08-11.
lostControllerState decides the three arms by the module's protocol.
simulation keeps Offline. modbus, opcua and ethernetip return
Unknown.
The consequence is not the status column. Offline is in the auto-HOLD set, so
before the split, powering off a controller node issued Hold to every Running
batch-allocated Unit reading a network module pinned to it. That included a Unit
which had just failed over successfully onto a healthy chassis under
ADR 0006 and whose I/O was
answering the whole time. Failover working correctly could hold the batch it had
just saved. Unknown is outside the auto-HOLD set, so the false trip closes with
the mapping.
Annunciation is unchanged. The alarm generator already raised the same
AlarmTypeSystem alarm at the same Medium severity for Unknown as for
Offline, so the split trades a wrong statement for a correct one and gives up
no signal.
The alternative this ADR left open — corroborating with the unit runtime, which
holds the refuting evidence — stays open and stays out of scope. It is not
needed to make the state honest. The runtime's own per-module verdict is served
by GET /api/v1/sites/{site}/units/{name}/runtime/diagnostics, and putting it
beside the module's state where an engineer can see the two disagree is
#1717.