Skip to content

ADR 0041: Redundant-collector sourcing answers the controller-failover gap; hot-standby FB-state replication is deferred behind a written requirement

Status: Accepted Date: 2026-08-06 Issue: #1018 (founder ruling 2026-08-06; supersedes the deferred half of #962)

Context

ADR 0026 settled the near-term half of #962 — bonded interfaces as the standard controller network posture — and deliberately deferred the harder question: what to build for continuous control and the data record across a controller death. Across that failure, a continuous control-module PID loop stops and re-establishes on the designated standby. The integrator lives in process memory (pkg/fbruntime/blocks/pid.go), runtime state persists to a hostPath that does not follow the pod (#731), and the loop restarts from configured initial conditions. The result is a bounded control gap — outputs hold last value, go to device fail-safe (ADR 0009), or are sequenced by an armed SafeStateChart (ADR 0008) — and a bounded data gap, because no reader is on the field during the RTO window. The gap is bounded, timestamped, and cause-attributed, and it is published in docs/ha-failure-modes.md rather than implied away.

#1018 carried the deferred decision between two candidate answers:

  • A. Hot-standby FB-state replication — bumpless continuous takeover, DeltaV parity. Expensive in a specific way: a standby holding live FB state interacts with the lease/epoch fencing that provides the one-writer guarantee, and it is precisely what turns a reverse-path partition into a dual-writer incident.
  • B. Redundant-collector sourcing — a redundant PLC or OPC UA server with historical buffering, backfilled by the historian on reconnect. It keeps the record whole even when control briefly holds, and it is a deployment-topology choice rather than product code.

The decision was gated on two inputs. The Strategy input landed: cndcs-business#30 was answered on 2026-08-04 (decisions/zero-gap-control-requirement.md), and its recommendation is exactly B. Its findings carry the weight here:

  • No regulation requires control-loop continuity. The binding texts are about records — 21 CFR 211.188, 211.68(b), Part 11 §11.10(b), Annex 11. ALCOA+ Consistent forbids unexplained gaps, and an explained gap is contemplated by the principle itself. The obligation is therefore to make the gap self-describing inside the batch record, not to eliminate it. A gap attributed only in a system log a QA reviewer never opens is still unexplained as far as the batch record is concerned.
  • The decisive finding is architectural rather than numerical. For every prospect in the August sequence, we are the orchestration layer above OEM control islands. The PID loops live in the skid vendors' controllers, and we do not replace them, so a failover in our tier cannot stop a loop we were never closing. That holds at any gap magnitude.
  • "Contractual" arrives through the URS chain, not through a regulator. A URS clause is a document we can ask for, and the discovery sheet now asks for it directly.

The second gate — measured bench numbers from #942 drills 9 and 10 — remains open behind procurement, and the ruling holds that it no longer gates. A bench number would matter only if the decision turned on how large the gap measures. The architectural finding makes the answer the same at any magnitude, so waiting for it is delay rather than caution. B is also cheap and reversible, which is exactly the kind of choice that tolerates deciding before the last input lands.

Decision

Redundant-collector sourcing is the answer to continuity across a controller failure. Hot-standby FB-state replication is deferred — not rejected forever, but parked behind named reversal triggers. Founder ruling on #1018, 2026-08-06. This supersedes the deferred half of #962 and completes the open item ADR 0026 § Decision (5) left behind.

  1. Redundant-collector sourcing is the recommended topology for an unbroken record. A redundant PLC or OPC UA server with historical buffering covers the data gap; the historian backfills on reconnect. It is deployment topology, not product code — the product ships the guidance, and no CRD field or controller grows from this decision.

  2. Hot-standby FB-state replication stays deferred. ADR 0006's rejection stands. The expensive option would trade a proven safety property — the one-writer guarantee that lease/epoch fencing provides — for a capability no committed prospect requires, and that trade is bad independent of the engineering cost.

  3. The reversal triggers are explicit, adopted from Strategy's answer:

  4. Any URS or RFP specifying controller redundancy. That is the mechanism by which a DeltaV capability becomes an "industry requirement" with no regulator involved. It is negotiable at requirements review and ruinous at OQ, and a URS clause is a document we can ask for.
  5. A continuous-manufacturing prospect (Continuus, Bright Path, Phlow, ODP) entering a real conversation. That is the only segment where we would close the loops ourselves and where the question is live.
  6. #942 reporting a gap materially larger than ADR 0026 assumes.

  7. The actionable engineering falls on the record, not the loop. Per the ALCOA+ reading, the failover gap must be self-describing inside the batch record rather than explained only in a system log. That work is #1308, and it is separate from the collector topology.

  8. Positioning: never claim bumpless parity. Redundancy is the incumbent's chosen ground, and asserting equivalence invites the one comparison we lose. We keep publishing the quantified gap in docs/ha-failure-modes.md, we quantify it further once #942 unblocks, and we put it in the diligence pack deliberately. In a room whose job is data integrity, a stated limitation is a better artifact than an unfalsifiable claim about a rare event.

  9. The IEC 62443 FR 7 availability targets stay deferred and ride with #942, not with this decision. ADR 0026 parked them until measured numbers exist, and nothing here settles them — the next reader should not treat them as revised by this ADR.

Alternatives Considered

  • Build hot-standby FB-state replication now (Option A). Rejected on three grounds. No committed prospect can require it, because we do not close their loops. It trades the one-writer safety property for bumpless takeover, which is a bad trade at any price. And its cost is continuous — state sync plus a correctness argument against the fencing model — where B's cost is a topology diagram.
  • Wait for the #942 bench numbers before deciding. Rejected by the ruling itself. The drills sit behind bench procurement, and the number they would produce cannot change an answer that rests on which loops we close. Drill 10 stays valuable as verification and as a diligence artifact, no longer as a gate.
  • Decide B but market it as redundancy parity. Rejected as positioning. DeltaV has shipped redundant pairs for twenty years, and the buyer's automation lead has run them. The quantified, cause-attributed gap is the stronger artifact in a data-integrity room.
  • Treat the published gap as sufficient and do nothing further. Rejected. ALCOA+ Consistent is satisfied by an explained gap, and today the explanation lives in the audit trail rather than in the batch record a QA reviewer actually reads. Closing that distance is the concrete obligation this decision produces (#1308).

Consequences

  • #1308 is the follow-up obligation: make the failover gap self-describing inside the batch record, bracketed the way ADR 0026 describes — a Hold event, a failover AuditRecord, and an explicit Restart — but carried where the QA reviewer looks. That is engineering work in the product, unlike the collector topology itself. #1753 carried the same obligation to the failovers the control lease does not drive, which is a planned-maintenance one and every failover of a unit in the default Autonomy mode. Those produce no lease records at all, so the gap is bounded from the re-bind and the re-bound runtime's return to normal, with the conservative opening bound disclosed.
  • An explained gap is explained only where the bracket describes the hole (#1805). The obligation is satisfied by the interval the record names being the interval the data is missing from, and the first implementation named one that was not. The data gap was bracketed by the control lease's own edges. The lease is an instrument for the one-writer guarantee and not for the record: it expires a whole lease duration after the runtime stops publishing, and the replacement takes it after that replacement has already resumed. Measured on the bench (#942 drill 10, n = 3) the record was missing 45.2 s of samples and the published bracket named 25.3 s of it, opening 28.0 s late and closing 8.0 s long. A reviewer looking inside that bracket found data, and more than half the missing samples fell outside it. The data gap now opens at the last exchange with the failed runtime the control plane had acknowledged. Each gap also declares whether its bounds are the gap's own edges or an interval containing it. Both of the data gap's bounds err outward, which is the direction that keeps the hole inside the explanation.
  • A bound kind is a claim about the data, so it is decided per branch (#1822). The sentence above is true of the acknowledged-contact opening and of nothing else, and the first implementation stamped Containing one line above the switch that chooses between three openings. Two of the three are fallbacks the assembler itself establishes are late, and each appends a PartialData warning saying so, while the published bound kind still promised the hole was inside the interval. Measured on the bench three months later, one rep of the same drill published Containing on a bracket that opened 94.5 s inside a 115.3 s hole and explained 103 of its 575 missing samples (#942 drill 10 rep 1, 2026-08-25). The fallback itself was correct and stays: that rep's operator instance started 18 s after the failed runtime stopped publishing, so it had never seen that runtime alive and had no contact to report, which is the ordinary shape of the first failover after a control-plane disruption. What changes is the label. A third value, Overlapping, names an interval whose closing bound errs outward and whose opening bound lies inside the hole. Two things follow. The enum stays a statement about where the data is rather than about which instrument produced the bound, so both fallbacks take the same value and the PartialData condition carries the difference between them. And a disclosure in a second field is not a substitute for a correct one in the first, because a reviewer who trusts bounds is never shown it.
  • docs/ha-failure-modes.md keeps the limitation published and gains the bench numbers when #942 lands. The redundant-collector guidance there is now the decided posture rather than a suggestion raised during design.
  • The marketing claim boundary from ADR 0026 is unchanged: never "bumpless", never "zero-gap", never redundant-pair parity with DeltaV.
  • The FR 7 availability targets in docs/ha-failure-modes.md are not revised. They ride with #942, per ADR 0026 and restated here deliberately.
  • No product code changes and no CRD surface, so reversibility stays high. The deferred hot-standby question remains where the irreversible commitment lives, and the reversal triggers above are the conditions under which it reopens.
  • One caveat is inherited from Strategy and recorded here: zero discovery calls have been held, and no prospect has sent a URS. The finding rests on regulatory texts, the prospect dossier, and what our architecture touches. That settles which loops we close; it does not close the positioning question permanently. Strategy reports back after the first three discovery calls, or immediately if a URS clause turns up.