Skip to content

ADR 0037: A server alarm is interpreted on the console it annunciates to, and the HMI gains no infrastructure surface

Status: Accepted Date: 2026-08-05 Issue: #1240

Context

#1168 built the Servers surface for the engineering UI and put the HMI out of scope. ADR 0032 then raised server degradation as an ISA-18.2 alarm in every site namespace, and the fan-out carried it onto the HMI for free: the sidebar cue, the topbar badge, the alarm list, the alarm detail panel and the retained MQTT event all took server alarms with no new API surface and no new poll.

The two decisions were each correct and together they left a gap. A plant-floor operator is now annunciated a condition their console cannot explain. What reaches them reads:

Server plant-w2 is not Ready (status False, reason KubeletNotReady)

That is a kubectl sentence on the surface built for the person #1168 described as not speaking kubectl. The engineering UI holds the rest of the story, and the operator does not have the engineering UI.

An operator looking at that alarm has three questions.

  1. Is the control system underneath me degraded?
  2. Are my units affected?
  3. What am I supposed to do about it?

Some of the answer is already on their console. The ADR 0023 health indicator sits in the HMI top bar on every screen, backed by the read-tier system:health action that every role carries. A unit whose runtime pod stops being ready raises its own System alarm at High through the equipment alarm generator, in plant language, naming the unit. Feed freshness (ADR 0028) marks the values on a frozen display as stale. So the operator is not blind to the effects of a dead node. What they cannot do is connect the effects to the cause, or tell a server alarm that threatens their batch from one that does not.

ISA-18.2 frames the defect precisely. An annunciated alarm must carry a defined operator response, and the response must be one that operator can carry out. Annunciating a condition whose response is unstated teaches the operator that the category is noise, which is the same failure mode ADR 0032 invoked when it refused to raise a second alarm for the shrinking etcd margin. The missing surface is not a node table. The missing thing is the response.

Decision

The HMI gains no infrastructure surface. Server degradation is answered by the health indication already on every HMI screen, and by the alarm itself, which states its consequence, its per-site impact and its operator response before any engineering detail.

The HMI gains no server inventory, no node detail and no infrastructure action

The asset tree is an engineering surface (ADR 0033), and every action on it is an engineering action (ADR 0034). An operator escalates a hardware fault. Kubelet versions, volume bindings, cordon state and quorum arithmetic are inputs to a decision the operator does not make, and putting them on the HMI would furnish the appearance of a decision without the authority to take it.

The health indication gains a servers dimension

GET /api/v1/system/health now grades the cluster's servers alongside its core services and its unit runtimes, using the one node-health derivation #1227 already moved server-side, so the dot on the operator's screen and the alarm on their console are graded by the same rules. ADR 0032 named this as a reasonable separate improvement, and it is the whole of the "minimal system health readout" the HMI needs: one chip, already present, already read-tier, on every screen.

The rollup grades a lost etcd quorum as critical and any server that is not healthy as degraded. A single server leaving Ready is a redundant element failing, and the workload it carried either reschedules or annunciates through its unit's own rule. Losing quorum is the cluster-level event that stops state changes taking effect, which is why it alone is graded critical.

The chip stays a state indication. It carries no acknowledgement, no history and no external annunciation, exactly as ADR 0023 requires, so nothing here competes with the alarm tier or annunciates a condition twice.

Every server alarm states its consequence, its impact and its response

A server alarm message is assembled in a fixed order:

  1. Consequence. What has happened, in terms of what the plant loses. This leads because the alarm list truncates a long message and the first sentence is what the operator reads at a glance.
  2. Impact on this site. Which unit runtimes in this site namespace were running on that server. ADR 0032's fan-out already writes one Alarm per site, so each site's copy names that site's own units, and a site with none on the server says so. That answer is worth as much as the positive one: it is how an operator learns the alarm does not threaten their batch.
  3. Response. Escalate to whoever maintains the control system hardware, stated on every alarm, because escalation is the defined operator response and an operator console cannot correct a server fault.
  4. Engineering detail. The readiness status, the reason string and the quorum derivation, kept last so the engineer looking at the same alarm loses nothing.

Where the impact cannot be determined, the alarm says so. A message that reported no affected units because the query failed would be read as reassurance, which is the rule #1230 established for the pre-flight panel and it applies with more force to an alarm.

Who to call is deployment data, and the product does not invent it

The alarm names the class of responder. A site's escalation contact belongs to that site's alarm response procedure, which is customer data the product has no business guessing. A configured per-site contact remains available as a later addition if a deployment asks for one.

Alternatives Considered

Give the HMI a minimal system health page: the node list trimmed to health and nothing else. Rejected on three counts. It is the Servers view with columns removed, so it re-poses every scoping question #1168 answered and then diverges from it. It needs either a new route or the servers:read action, which a locked-down operator role may deliberately not carry. And it answers "which server" for a reader who has no use for the answer, while the question they can act on is whether their units are affected, which a node list does not address at all.

Write down that the operator escalates, and change nothing else. This is the option the issue offers as defensible, and it is half right. Escalation genuinely is the correct response, and the ruling above adopts it. It is wrong about where the writing goes. A response recorded only in a manual leaves the operator inferring it from a kubelet reason string at the moment they are holding an unacknowledged alarm, and ISA-18.2 asks for the response to travel with the alarm.

Suppress server alarms from the HMI and keep them engineering-only. Rejected. It hides a plant-relevant condition from the person watching the plant, which is the outcome ADR 0032 exists to prevent. A lost quorum also changes what the operator can do, because state changes stop taking effect, so it must reach them.

Put the server cause on the unit's own runtime alarm instead, and leave the server alarm as it is. Attractive, because it lands the infrastructure fact on the alarm the operator already acts on. Rejected as the primary mechanism because it covers only the case where a unit runtime is lost. A control-plane failure, a resource pressure warning and a lost quorum all reach the operator with no unit alarm beside them, and those are the cases where the operator most needs telling that the fault is not theirs to fix. It remains a reasonable addition on top of this decision.

Fold the affected-unit list into a structured field on the Alarm CRD rather than into the message. Rejected for now. It would need API surface, a DTO field, HMI rendering and a policy for every other alarm source that has no such list. The message reaches every consumer that already carries alarms, including the retained MQTT event an external annunciator reads, at no schema cost.

Consequences

  • The health chip on every HMI screen now moves for a hardware condition. A plant running a chronically pressured node will see a persistently degraded chip. That reading is true, and the remedy is to fix the node.
  • Server alarm messages get materially longer. The alarm list truncates the row and carries the full text in the tooltip and the detail panel, so the leading consequence sentence is load-bearing and the rules table in ADR 0032 now governs severity alone.
  • ServerAlarmReconciler reads Unit objects in each site namespace and watches for runtime-binding changes, so a failover updates the impact sentence on a standing alarm rather than leaving a stale claim on it.
  • A cluster where the reconciler cannot list units keeps annunciating, with the impact stated as undetermined.
  • The system health summary names a degraded server to every read-tier role. Nothing is newly exposed: ADR 0032 already puts the server's name on an alarm in front of every operator on every site.
  • Server health now has one derivation with three readers, which are the Servers view, the health rollup and the pre-flight panel. The alarm reconciler keeps its own evaluation because it grades ISA-18.2 severities rather than a health label, and both are stated against the same rules table.

Amendment: a server held out of service is not a healthy server (#1859)

The decision above grades the servers dimension off one derivation, serverNodeHealth, and says any server that is not healthy is degraded. That derivation reads the Ready condition and the three pressure conditions, and it is explicit that cordoning is deliberately absent from it. A cordoned machine has nothing wrong with it.

A node the product itself has taken out of service is therefore graded healthy for the whole time it is out of service. A cordon and a drain both leave the kubelet Ready and raise no pressure condition, so nothing this rollup read could see either one.

Measured on the bench. cp-1 sat in a NodeMaintenance at phase Draining for eighteen hours, with site-bench-site/server-cp-1-maintenance-overrun unacknowledged for the last seven of them, and GET /api/v1/system/health called the cluster's servers healthy throughout. The one thing it reported was the broker, degraded because the drain on cp-1 was evicting it once a minute. It named the symptom and had no way to name the cause standing one object away. #1851 ruled out the ReplicaSet, a rollout, Helm, Flux, taint eviction, node pressure, a container fault and the pod reaper before the maintenance came into view.

A server is graded on two questions

Is the machine working, and is it carrying work. serverNodeHealth answers the first and is unchanged, so the health column on the Servers page still means what #1227 moved that derivation server-side to make it mean. outOfServiceReason (internal/gateway/diagnostics.go) answers the second, from spec.unschedulable and the open NodeMaintenance, and only this rollup reads it.

The two answers land on one row per server. A shutdown maintenance leaves the machine unreachable, the readiness derivation grades it offline for exactly the right reason, and the maintenance is the sentence an operator needs beside it.

The entry is as loud as the fact, and no louder

Degraded, never critical. The machine is running, quorum is unaffected, and what is lost is capacity somebody asked for. A lost quorum stays the one server condition graded critical, on the reading the decision above already gives.

Silence needs a reason, and the reason is a window the product owns

An open maintenance inside its ADR 0036 alarm-shelve window is a hold the product asked for. It carries a requester, a justification and an expected end, and the alarm tier is deliberately quiet for that window about the alarms the maintenance is expected to cause. The rollup is quiet on the same clock, taken from the same call, NodeMaintenance.AlarmShelveDeadline. Past that instant server-<node>-maintenance-overrun fires, and this entry now appears with it. Before this amendment that alarm stood on the operator's console while the health chip on the same screen read green.

A cordon no maintenance of ours accounts for gets no such grace. Nothing in the product asked for it and nothing in the product will return it. Both maintenance verbs refuse one outright, with the message that it has to be returned to service the same way it was taken out, so there is no instant at which it is expected to end. A hold the product can neither explain nor finish is the last thing that should read as healthy.

The second read has its own unanswered case

The out-of-service question is answered from a second object. A NodeMaintenance list the gateway could not read therefore leaves it unanswered for every server at once, and that is reported as its own entry beside the unreadable-node-list one #1103 added. Passing it as a cluster holding nothing out of service would be this amendment's own defect, one layer down.

Consequences of the amendment

  • A deployment that leaves a node cordoned reads degraded until it is uncordoned. That reading is true, and the remedy is to return the node to service or to record a maintenance for it.
  • The health chip and the maintenance-overrun alarm move together. They are graded from one derivation of the window, so a future change to MaintenanceAlarmShelveWindow moves both.
  • Server health still has one derivation with three readers. The out-of-service question is a second derivation with one reader, deliberately, because the Servers page renders cordon state as its own column and grading its health column on a cordon would paint a chassis with nothing wrong with it.