Network Requirements¶
This page is the single consolidated reference for a plant IT team standing up Cloud-Native DCS on an existing OT/IT network. Use it for firewall change requests, subnet design, DNS and NTP planning, and switch / VLAN configuration.
For the security rationale behind each rule, see Security Hardening and Threat Model. For the underlying port allocation policy, see Architecture § Port Allocation Policy.
Zone Model¶
Cloud-Native DCS follows the IEC 62443 zone/conduit model. A typical deployment uses four zones:
| Zone | Contents | Reachability |
|---|---|---|
| External | Users' browsers, CLI workstations, SIEM collector, monitoring (Prometheus, Grafana) | Reaches the Control Zone through a single ingress (HTTPS 443) |
| Control | Kubernetes control plane, gateway, operators, MQTT broker, historian | Reaches Field Zone on mTLS 8883 (MQTT) and the runtime HTTP API. Reaches External Zone only via ingress, unless the optional northbound OMF egress, the optional Object Lock audit mirror, or the optional database backup to an external object store is enabled — see Control → External. |
| Field | Edge devices running unit-runtime pods (x86_64 industrial PCs on an OT VLAN) |
Reaches Field Bus Zone on industrial protocols. Reaches Control Zone on mTLS 8883 (MQTT) only. |
| Field Bus | Physical I/O modules (Wago 750, Turck TBEN, remote I/O couplers, field instruments) | Not reachable from anywhere except its own Field Zone runtime. |
Physical-Layer Posture (Controller Nodes)¶
The zone model above is logical. This section is the physical counterpart, and it is a requirement for controller (edge) nodes, because a controller with a single field-network link treats a pulled cable exactly like a dead node.
Bond every controller network zone, active-backup. Each zone gets its own bond of two physical ports. Where the deployment has redundant switches, land the two members on different switches.
What a second switch does and does not buy
Separate switches do not make a field-network switch failure survivable on their own. Remote I/O is single-homed: a coupler, a drop, or a PLC has one port on one switch. When that switch dies, the controller's surviving bond member has nothing left to reach. Splitting I/O across two switches turns a total field outage into a partial one. That is useful, but it is not redundancy. Genuine field redundancy needs a ring topology (MRP/RSTP/DLR) or dual-homed field devices, which is a larger design than bonding.
The zone where two switches genuinely pay off is the supervisory one: the far end is the multi-node control plane and can be split across both switches. Prioritize there when ports are scarce.
Bonding is worth doing on a single switch regardless. The failure it eliminates is a cable, NIC, or switch-port fault on a live controller, and that needs no second switch.
| Zone on the controller | Ports | Posture |
|---|---|---|
| Field (L1, to remote I/O) | 2 | Bonded — required. Losing this link stops regulation. |
| Supervisory (L2, to the API server) | 1 or 2 | Bonded if the port budget allows. A supervisory outage costs control-plane reachability only, which Autonomy mode already rides out. |
That is a 3-port minimum for a controller node, 4 ports for the full posture. Size the hardware accordingly. A dual-NIC industrial PC cannot carry this, because both of its ports are already spent on zone separation.
Do not trunk both zones over one bond. Carrying supervisory and field VLANs on shared physical media buys link redundancy at the cost of the segmentation the IEC 62443 posture depends on (SR 5.1 RE(1)). Two ports is not enough hardware for this design. Use four.
Choose active-backup over LACP. The bandwidth envelope below is nowhere near a 1 GbE link, so aggregation buys nothing we need. Active-backup requires no switch-side configuration and spans two independent switches without MLAG or stacking. Deployments already running MLAG throughout may use LACP as a local choice. The guidance stays active-backup.
Bonding is configured in the node's machine config at deployment time
(Talos exposes a native bond interface type), consistent with node join
being a deployment-layer concern. It is transparent to the product: a
bond presents one logical interface, so the function-block scan, the
protocol drivers, and hostNetwork: true port binding are all unaware of
it. No DCS configuration references a bond.
Rationale and the rejected alternatives are in ADR 0026. What bonding does not cover (the control and data gap across an actual controller death) is documented in HA and Failure Modes.
Required Ports¶
External → Control¶
| Port | Protocol | Purpose | Notes |
|---|---|---|---|
| 443 | HTTPS | Gateway REST API + Web UI | TLS-terminated at ingress; OIDC Bearer auth |
| 443 | HTTPS (WebSocket upgrade) | Live HMI updates | Same port as REST |
No other external-to-control ports should be open.
Control → External (northbound)¶
Three optional components send data out of the cluster to a host the customer operates. Everything else in the product reaches only what the chart deployed, so a plant that has enabled none of them needs no rule here at all.
| Port | Protocol | Purpose | Notes |
|---|---|---|---|
| 443 | HTTPS | omf-egress → the site's OMF endpoint (PI Web API, Edge Data Store, AVEVA Data Hub) |
Outbound only. Off by default (omfEgress.enabled: false); the port follows omfEgress.egress.ports |
| 443 | HTTPS | audit-archiver → the Object Lock bucket that holds the immutable archive copy |
Outbound only, hourly and only when a batch crosses the retention cutoff. Off by default (historian.audit.archival.immutable.enabled: false); the port follows historian.audit.archival.immutable.egress.ports. Needs no rule when the bucket is served inside the cluster |
| 443 | HTTPS | historian database (CNPG) → the S3-compatible store holding its WAL segments and base backups | Outbound only, continuously while the database is running. Off by default (historian.backup.enabled: false); the port follows historian.backup.s3.endpointURL, so an in-cluster minio on :9000 needs no rule here |
The direction is worth stating plainly in the change request, because a historian feed sounds bidirectional and none of these is. Nothing is read back over the OMF rule, and the PI System cannot write into the DCS through it. The other two carry traffic in one direction too, though what they write does get read back. The audit mirror is read by hand during a restore, from an operator's workstation outside the cluster, and the database backup by CNPG itself when a cluster is restored from it.
The in-cluster half of the same statement is a NetworkPolicy, and it needs the
endpoint written as CIDRs (a NetworkPolicy resolves no hostname):
omfEgress.egress.destinationCIDRs,
historian.audit.archival.immutable.egress
and
historian.backup.s3.egress.
A firewall rule opened without the matching chart value produces a component
that still cannot reach the endpoint, and none of them reports that as a
policy denial: the OMF egress reports an endpoint outage, the archiver's
mirror upload fails on the next run that has something to archive, and the
database keeps its unarchived WAL on the data volume until the volume fills.
Control → Field¶
| Port | Protocol | Purpose |
|---|---|---|
| 8883 | MQTTS | Telemetry, commands, state transitions, alarms |
| 61152 | HTTPS + mTLS | Control-operator → unit-runtime HTTP API (FB network deploy, read/write) |
Unit runtime pods use hostNetwork: true, so these ports land on the
node IP. No ClusterIP is involved.
Field → Field Bus¶
| Port | Protocol | Purpose |
|---|---|---|
| 502 | Modbus TCP | Wago 750-series and other Modbus TCP couplers |
| 44818 | EtherNet/IP | EtherNet/IP remote I/O (explicit messaging) |
| 2222 | EtherNet/IP (I/O) | EtherNet/IP cyclic I/O messaging (UDP) |
| 4840 | OPC UA | OPC UA servers on PLCs and gateways |
Intra-Control (cluster-internal)¶
The following ports are only relevant if NetworkPolicies are being authored by hand. The shipped NetworkPolicies handle them automatically.
| Port | Protocol | Component |
|---|---|---|
| 8090 | HTTP | Gateway internal |
| 8092 | HTTP | Historian |
| 8094 | HTTP | Historian metrics (plaintext, in-cluster only) |
| 8095 | HTTP | OMF egress health (plaintext, in-cluster only) |
| 8096 | HTTP | OMF egress metrics (plaintext, in-cluster only) |
| 8080, 8081 | HTTP | Operator metrics / health |
| 8090 / 8091 | HTTP | io-probe (per-Controller) |
See Architecture § Port Allocation Policy
for the full table and the reason ports avoid the Linux ephemeral port
range 32768–60999.
Firewall Change Request Template¶
Paste this into your firewall change request, filling in subnets:
ALLOW External-user-subnet → Control-ingress-VIP tcp/443
ALLOW Control-subnet → Field-subnet tcp/8883
ALLOW Control-subnet → Field-subnet tcp/61152
ALLOW Field-subnet → Field-bus-subnet tcp/502
ALLOW Field-subnet → Field-bus-subnet tcp/44818, udp/2222
ALLOW Field-subnet → Field-bus-subnet tcp/4840
ALLOW Control-subnet → NTP-server udp/123
ALLOW Field-bus-subnet → Field-subnet udp/123
ALLOW Control-subnet → DNS-server udp/53
DENY Field-bus-subnet → (any) (default)
DENY Field-subnet → External-subnet (default)
Add this line only when the northbound OMF egress is enabled, and delete it again if the plant later turns it off:
ALLOW Control-subnet → PI-Web-API-host tcp/443
DNS¶
| Name | Resolves to | Used by |
|---|---|---|
dcs.<plant>.example.com |
Control-Zone ingress VIP | External-Zone clients (browsers, CLI) |
docs.<plant>.example.com |
Control-Zone ingress VIP | External-Zone browsers, when docs.ingress gives the documentation site its own name |
| Gateway mTLS SNI name | Gateway pod IP (cluster-internal DNS) | Operators, runtime pods |
| MQTT broker name | Broker service IP (cluster-internal DNS) | All MQTT clients |
cert-manager issues certs pinned to these names. If you change a DNS name, you must re-issue the corresponding Certificate CRs. Pinning certs to IPs means the same IP changes break the trust chain. Always use DNS names.
Time Synchronisation¶
21 CFR Part 11 §11.10(e) requires "secure, computer-generated, time-stamped audit trails". Clock discipline is a prerequisite for a usable audit trail and for correct batch-record correlation across units. Target:
- Stratum 2 or better NTP source reachable from every node in the Control Zone and Field Zone.
- Chrony or systemd-timesyncd enabled and in sync on every node.
chronyc trackingshould report an offset under 50 ms. - No wall-clock skew larger than 1 second between any two nodes in the cluster.
PTP (IEEE 1588) is not required. Scan rates are single-digit to hundreds of milliseconds, which NTP comfortably covers.
Field devices need a time source too, and usually cannot reach yours¶
The targets above are scoped to the Control Zone and Field Zone. Devices in
the Field-Bus Zone (couplers, PLCs, remote I/O) are not covered by them and
usually cannot be, because the firewall template above ends with
DENY Field-bus-subnet → (any). That denies udp/123 along with everything
else, so a field-bus device has no path to the NTP source the rest of the
network uses.
This matters more than an unset clock on a node. A field device with a wrong clock stamps certificates from it, and a certificate carries the bad time forward until someone regenerates it. A self-signed device certificate with a 30-day validity window, issued from a clock several years in the past, stays silently expired long after the clock is corrected. The device will not flag it, because from its point of view the certificate was valid when issued.
Three options, in the order they are usually preferred:
| Option | When it fits |
|---|---|
| NTP served from the hosts that already bridge both zones | The nodes running unit-runtime are dual-homed into the Field Zone and the Field-Bus Zone by definition. Serving udp/123 on their field-bus-facing address adds no route and no gateway; the device initiates to the node. Usually the cheapest correct answer. |
| A dedicated NTP appliance inside the Field-Bus Zone | The field segment must hold time with the control system down, or the site wants the zone genuinely standalone. Costs hardware and a battery-backed RTC. |
| Manual clock-set with a documented review interval | Only where the device has no NTP/SNTP client at all. The interval must be shorter than the shortest certificate lifetime on the device — commonly 30 days for a self-signed default, which makes this expensive to sustain. |
If time is served from the bridging nodes, declare the conduit rather than letting it exist informally:
ALLOW Field-bus-subnet → Field-subnet udp/123
Note the direction: the field-bus device is the client. Nothing is forwarded between zones and the field-bus subnet still needs no default gateway.
Stratum expectations differ by zone. A local reference on an isolated segment serves a high stratum, commonly 10, because it is not traceable to a reference clock. That is the honest number and it is fine. The Field-Bus Zone requires that certificates carry sane validity windows and that event timestamps correlate. Traceability is a separate requirement belonging to a different zone. Keep the stratum-2 target for the Control Zone, and do not configure a device to reject a high-stratum source that is the only one it can reach.
Commissioning order¶
Set the clock and provision the time source before generating any certificate or enabling any security feature on a field device. Doing it afterwards means regenerating every certificate issued in the meantime, and the ones that get missed fail silently later.
For each field device, in this order:
- Set date, time and timezone. Timezone matters independently of the instant: a device left on a factory-default zone records local timestamps that do not correspond to the plant's, which corrupts event correlation even when UTC is right. Prefer UTC on devices whose logs are read alongside cluster logs.
- Point the device at the segment's time source and confirm it has actually synchronised by reading the clock back out of band. A device web UI that reports success is weak evidence, and so is one that reports failure.
- Only then generate certificates or enable secured endpoints.
- Record the device's clock behaviour in the commissioning record, including whether it holds time across a power cycle. A device with no RTC battery returns to its factory clock on every cold start, so its time source must already be reachable while it boots.
If a device was commissioned in the wrong order, regenerate every certificate it holds after fixing the clock, and check the validity windows rather than assuming.
Bandwidth Envelope¶
The shipped telemetry pipeline is lightweight:
| Load | Typical bandwidth |
|---|---|
| Per unit runtime — idle | <10 kbps |
| Per unit runtime — 1 batch running, 20 tags at 200 ms | ~50–100 kbps |
| Gateway REST API — per active operator | <50 kbps |
| Historian ingest — 1 site, 10 units | ~1 Mbps sustained |
A 10 Mbit/s link between Control and Field zones is ample for any single site with tens of units. A 100 Mbit/s link handles any realistic pharma batch plant. Latency and packet loss are usually the binding constraints, well before bandwidth.
QoS and Latency Tolerance¶
| Path | Target round-trip | Hard ceiling |
|---|---|---|
| Gateway ↔ runtime HTTP API (mTLS 61152) | <10 ms | 100 ms — above this, hot-swap and health checks start timing out |
| MQTT publish → subscribe (8883) | <100 ms | 2 s — above this, historian gaps and UI live-update lag |
| Runtime ↔ field I/O (Modbus, EIP, OPC UA) | <20 ms | 200 ms — above this, scan-cycle budget is blown |
The store-and-forward queue on the runtime tolerates short MQTT broker outages without telemetry loss. There is no equivalent for runtime ↔ I/O outages. Those manifest as driver-health alarms within one scan cycle.
What Is NOT Required¶
- Windows engineering station. Authoring, operation, and review all happen in the browser. There is no native Windows client.
- Dedicated OT DNS. Cluster-internal DNS (CoreDNS) handles intra-cluster resolution. Plant-wide DNS only needs the external gateway name.
- Active Directory. OIDC is the supported authentication protocol, and AD/Entra integrates through it. (LDAP bind is supported as a fallback, see Security Operations § LDAP Authentication.)
- Windows file shares. No SMB, no CIFS, no NetBIOS traffic anywhere.
- Printers. Batch record export is PDF/JSON via the browser or API. Printers are the customer's own infrastructure and stay outside the DCS.
Related Documentation¶
- Security Hardening — the production configuration checklist that turns the zone model into NetworkPolicy YAML.
- Device Enrollment — how a new edge device joins the Field Zone.
- HA and Failure Modes — what a link fault, a partition, and a controller death each do to control and to the record.
- Reference Architectures — the hardware patterns these requirements assume.
- Deploy Your Own — the reference single-site install walkthrough.