Skip to content

Upgrade and Rollback

Status: Procedures below match the current Helm chart and the physical-operator's runtime-pod behaviour. The version compatibility matrix is a placeholder -- Cloud-Native DCS does not yet have stable release tags to map against, so that table lights up once v0.x releases exist.

Procedure for upgrading Cloud-Native DCS in place on an existing Kubernetes cluster and rolling back when an upgrade fails verification. Platform-specific steps (etcd snapshots) show both Talos (the reference deployments) and k3s forms.

Before You Start

  • Read the release notes for the target version. Every release lists breaking changes, CRD schema migrations, and expected downtime.
  • Confirm a current backup exists -- see Backup and Recovery. At minimum:
  • etcd snapshot from the last hour: talosctl -n <node> etcd snapshot pre-upgrade.snapshot (Talos) or k3s etcd-snapshot save (k3s)
  • CRD export: dcs backup crds -o /backups/pre-upgrade-$(date +%F).yaml
  • HMAC signing key Secret: kubectl get secret dcs-signing-key -n kube-system -o yaml > /backups/dcs-signing-key.yaml
  • Verify the current deployment is healthy. Upgrading on top of a degraded system multiplies the failure modes you have to debug:
    dcs health
    kubectl get pods -n dcs-system
    kubectl get certificate -n dcs-system -l app.kubernetes.io/part-of=cloud-native-dcs
    
  • Schedule the window. Even rolling upgrades briefly interrupt per-unit runtime pods (see Downtime Expectations). Plan for a quiet period when no batch is in Running or Holding on a unit you intend to roll.

Architecture Facts That Matter for Upgrades

Two chart behaviours tend to surprise first-time upgraders:

  1. Helm installs CRDs but never upgrades them. The chart ships the CRDs in its crds/ directory, so a fresh helm install creates them. Helm deliberately skips the crds/ directory on helm upgrade, and helm rollback will not undo a CRD change either. On every upgrade you must apply the new CRDs yourself with kubectl apply (or make install, from config/crd/bases/) before running helm upgrade.
  2. Unit runtime pods and io-probe pods are reconciled by the operators alone. Helm never touches them. They run as bare Pods in site-<name> namespaces, so the chart upgrade rolls every Deployment and leaves these two carrying whatever image the operator that created them chose.

The two halves behave differently, and the difference is deliberate.

Unit runtime pods roll themselves where the plant allows it (#1762). The physical-operator compares the running image against the one it would create the pod with today. A unit is free when it has no allocated batch, or when its batch is in a resting phase. For a free unit the operator recreates the pod on the current image during its next reconcile, and writes an AuditRecord naming both images. Where a batch is actively driving I/O, it changes nothing and reports instead. status.conditions[type=RuntimeImageDrift] goes True with reason BlockedByBatch, naming the batch, and dcs_unit_runtime_image_drift reads 1. Pulling the runtime out from under a Phase mid-step can leave actuators in indeterminate states, which is the same reason dcs runtime restart refuses it.

The runtime roll is therefore a step you verify. What remains for you is the units a batch was running on, and step 5 below is how you find those and account for them.

io-probe pods do not roll themselves. Nothing recreates a Running probe pod on a superseded image, and that stays true on purpose: its read cadence feeds every device comm-loss watchdog a module declares through spec.failSafe, so an unattended roll is a plant action. The operator reports the drift on every affected IOModule as an IOProbeImageDrift condition, and dcs io probe restart is the audited remedy. Rolling probes is step 6, and it is a deliberate gesture.

One-time migration: per-site runtime Certificate (v0.x, issue #213)

Upgrading from a chart version that predates #213 switches the CA issuer from namespace-scoped Issuer to cluster-scoped ClusterIssuer and moves per-site cloud-native-dcs-runtime-mtls Secret provisioning from a gateway-managed copy to a cert-manager Certificate owned by the Site reconciler. After helm upgrade:

# 1. Verify cert-manager is configured with --cluster-resource-namespace
#    pointing at the DCS release namespace (usually dcs-system). Without
#    this, cert-manager cannot find the CA Secret and ClusterIssuer stays
#    NotReady.
kubectl get clusterissuer cloud-native-dcs-mtls-ca

# 2. Delete manually-copied runtime Secrets in each site namespace so the
#    Site reconciler can recreate the Certificate → Secret chain cleanly:
for NS in $(kubectl get ns -l dcs.io/site -o name); do
  kubectl delete secret -n "${NS#namespace/}" cloud-native-dcs-runtime-mtls --ignore-not-found
done

# 3. Trigger a Site reconcile so the reconciler provisions a Certificate.
#    Any metadata change works — bump an annotation:
#      kubectl annotate site <name> \
#        dcs.io/reconcile-requested-at="$(date -u +%FT%TZ)" --overwrite
#    cert-manager then issues a fresh Secret within seconds.

# 4. Bounce runtime pods once so they mount the fresh cert content:
kubectl get pod -A -l dcs.io/component=unit-runtime -o name | \
  xargs -I {} kubectl delete {} -n <namespace>

Set mtls.certManager.clusterIssuer=false and mtls.certManager.perSiteRuntimeCert=false if your cert-manager install can't be reconfigured. The pre-#213 behaviour is preserved, but you keep the manual-copy drift risk.

Archive-integrity scheduler enabled by default (issue #216)

After upgrading to a chart that includes the Phase B scheduler, the gateway runs dcs audit verify --archived on gateway.archiveIntegrity.interval (default 6h) and writes one AuditRecord per run. Set gateway.archiveIntegrity.enabled: false or env GATEWAY_ARCHIVE_INTEGRITY_ENABLED=false to keep the old quiet behaviour on downgrade-friendly test clusters. No schema change is required. The scheduler reads the same audit_archive_manifest rows written by the archiver in Phase A.

Historian disk-watchdog removed (issue #243)

The chart-shipped historian-disk-watchdog CronJob (the interim coverage for #237/#238) has been removed. The two alerts it implemented are now opt-in PrometheusRule resources gated on monitoring.prometheusRule.historianDisk.enabled (default false). On upgrade:

  • Any value set under historian.alerts.watchdog.* becomes a no-op. Helm silently ignores unknown keys, so the upgrade itself does not fail. The CronJob is removed from the cluster as part of the rollout.
  • If you rely on the watchdog's webhook, install kube-prometheus-stack (or any Prometheus stack) and turn the new rules on:
    monitoring:
      prometheusRule:
        historianDisk:
          enabled: true
    
    The reference deployment's Flux overlay (cndcs-deploy-demo/flux/clusters/demo/) is the worked example.
  • The Slack/Discord/PagerDuty webhook URL is now consumed by Alertmanager, not the CronJob. Move the demo-alert-webhook Secret from dcs-system into the monitoring namespace, or recreate it there.

Production deployments should keep historianDisk.enabled: false and rely on TimescaleDB native retention plus a CSI-backed PVC instead. See production-deployment.md § 1, § 4.

Optional S3 Object-Lock mirror (issue #216, Phase C)

The audit-archiver chart now accepts historian.audit.archival.immutable.* values. It is off by default. Existing deployments keep the current PostgreSQL-only archive behaviour until an operator explicitly points the archiver at a pre-provisioned Object-Lock bucket plus a Secret with S3 credentials. Enabling mid- cluster is safe. No migration is required, because PG remains the primary query path and the mirror only covers batches written after the flip. Bucket Object Lock cannot be added to an existing bucket retroactively, so new buckets must be created with --object-lock-enabled-for-bucket (AWS) or mc mb --with-lock (minio). See backup-recovery.md for the full checklist.

The mirror now needs its egress destination written down (issue #1514). A deployment that already runs the mirror with networkPolicies.enabled will find the upgrade refused at render time until historian.audit.archival.immutable.egress.destinationCIDRs (or destinationPodLabels, for a bucket served inside the cluster) is set. The refusal names the value and an example. It is not a new requirement so much as an old one that was never expressible: the archiver's NetworkPolicy has always declared policyTypes: [Egress] and has never carried a rule for the mirror endpoint, so on any cluster whose CNI enforces NetworkPolicy the mirror upload was already being denied. The PostgreSQL archive kept working, and only the 21 CFR Part 11 §11.10(c) restore copy went missing. Nothing caught it because every stack the mirror had run on used a CNI that ignores NetworkPolicy entirely.

Two things to check on the way through, and they are checked in different places. Confirm the archiver's recent runs succeeded (kubectl get jobs -n dcs-system over the CronJob's history), because a run that had a batch to archive would have failed on the denied upload and left those AuditRecord CRs in etcd, waiting for a run that can complete. Then list the bucket itself. A period during which the archiver was denied is a period with no immutable copy, and no later run rewrites it: a batch is mirrored once, by the run that archives it. dcs audit verify --archived will not show that hole, because it verifies the manifest chain in PostgreSQL and never reads the bucket.

The historian database backup now needs its egress destination written down too (issue #1516). A deployment that runs historian.backup.enabled with networkPolicies.enabled will find the upgrade refused at render time until historian.backup.s3.egress.destinationCIDRs (or destinationPodLabels, for a store served inside the cluster) is set. It is the same class of gap as the mirror above, in the policy one file over: the CNPG pods' policy declares policyTypes: [Ingress, Egress] and its backup rule named no destination at all, allowing every address in the world on TCP/443, while the port the chart's own endpointURL example uses is 9000. So a store on any port but 443 was already being denied. The port now comes off endpointURL, and the hard-coded 443 is gone.

Check the store for a recent base backup on the way through, because a denial here does not look like one. PostgreSQL does not drop a WAL segment it has not archived, so pg_wal grows on the data volume until the volume fills. That fill is what the 2026-04-12 demo incident was, one cause upstream.

IOModule session credentials move to a Secret (issue #1912)

An OPC UA IOModule authored before this release carried username and password in spec.options. Those keys still open the session after the upgrade. Nothing stops. Each such module reports status.conditions[CredentialsSecured] as False/PlaintextInSpec, dcs io get prints Credentials: PLAINTEXT in spec.options, and the gateway stops serving the password on any read (ADR 0087).

Migrate a module by writing its credentials into a Secret in the site namespace and replacing the option keys with a spec.security block that names it (see Opening a secured session). The two are refused together at admission, so the migration is one write. The unit-runtime pod serving that module is recreated to mount the Secret, which is a runtime restart for that unit. The io-probe is not recreated. The module reports ProbeNotProjected until dcs io probe restart is run, and that is the same read-gap decision step 6 below describes.

Once every module is on a Secret, set ioSecurity.refusePlaintextCredentials: true in the chart values. From then on the driver refuses a module carrying the deprecated keys in both pods, which is the declared boundary the deployment asked for.

The IOModule CRD published with chart 0.7.3 does not install on a Kubernetes 1.31 apiserver: one of its validation rules spelled credentialsRef.namespace where CEL needs __namespace__, and the apiserver refuses the whole CRD with undefined field 'namespace'. A 1.33 or newer apiserver admits it. On an older cluster, take the next patch release, whose CRD carries the escaped spelling. Nothing else in 0.7.3 is affected, and a cluster that already holds the CRD keeps it.

The pod reaper CronJob moves to Replace (issue #1937)

Charts before 0.7.6 ran the terminated-pod reaper under concurrencyPolicy: Forbid. A reaper Job whose pod the scheduler refused stayed active with nothing to end it, and Forbid then suppressed every later run. Chart 0.6.3 added activeDeadlineSeconds to the job template, but that field lives on the Job, and a Job is the CronJob's child. An upgrade rewrites the CronJob and never touches a Job that already exists, so a cluster that was already parked stayed parked through every upgrade since. On the reference bench one Job created on 2026-08-25 survived five chart upgrades and nine days.

From 0.7.6 the CronJob runs under Replace. At the first :23 after the upgrade the controller deletes any Job still active from the old chart and creates the new run. There is nothing to run by hand, and no reaper run is lost that the old policy would have kept. Confirm it afterwards:

kubectl -n dcs-system get cronjob -l app.kubernetes.io/component=pod-reaper

ACTIVE reads 0 between runs and LAST SCHEDULE is under an hour old. The System Health page carries a Pod Reaper card from this release. It reads degraded whenever the reaper has not succeeded in two schedules, so the parked state is visible without reading the namespace by hand.

The archiver and prune CronJobs follow, and the prune alert reads silence (issue #1942)

The audit archiver and the historian prune CronJobs move to Replace in the same release, and each Job template gains a deadline sized just under its own schedule: 55 minutes for the hourly archiver (historian.audit.archival.activeDeadlineSeconds) and five hours for the six-hourly prune (historian.prune.activeDeadlineSeconds). A run parked on an older chart is freed at the first tick after the upgrade, the same way as the reaper. An archiver run killed at its deadline resumes on the next tick with nothing lost, because every batch is durable before its records leave etcd. A record the killed run had attested but not yet deleted is deleted by the next run. Charts before 0.7.6 skipped such a record forever.

The Historian Prune card joins the System Health page. It reads degraded when the prune has not succeeded in two of its six-hour schedules.

DCSHistorianPruneJobFailing no longer requires a failed Job. Charts before 0.7.6 keyed it on kube_job_status_failed. A Job whose pod never scheduled never fails, so the rule could not see the parked shape at all. It fires on silence now: no success in pruneStaleHours measured off the CronJob's own last success, its creation when it has never succeeded, or any labelled Job's completion. A suspended CronJob is exempt. The rule reads three series from the kube-state-metrics cronjob collector that the old one did not, kube_cronjob_status_last_successful_time, kube_cronjob_created and kube_cronjob_spec_suspend. A kube-state-metrics too old to export them leaves the rule reading the Job completions alone. A manual prune still moves those.

Version Compatibility Matrix

Cloud-Native DCS is pre-1.0 and does not yet have release tags to map against. This table will light up as v0.x releases ship.

From To Forward-compatible Rollback-safe Notes
pre-release pre-release n/a n/a No stable release yet

When releases begin: - Forward-compatible means the new operators tolerate the old CRD schemas for one minor version, so a rolling upgrade is safe. - Rollback-safe means the old operators can still read any CR that was written by the new version. A rollback-unsafe upgrade typically requires an etcd snapshot restore. A helm rollback cannot do it.

Kubernetes Version Upgrades

Upgrading the cluster underneath Cloud-Native DCS is a separate exercise from upgrading the chart, and every constraint on it comes from upstream. Cloud-Native DCS builds against the Kubernetes 1.36 API libraries and runs its envtest suites on the 1.36 control-plane binaries. The reference deployments run Talos Linux v1.10.9 with Kubernetes v1.33.6, and the documented floor for a customer cluster is Kubernetes 1.27.

The wave of upstream removals landing between Kubernetes 1.35 and 1.38 is about node prerequisites. None of it lands on the DCS control plane. Read the table before scheduling a cluster upgrade. Two of these rows fail the kubelet on a node, which is a worse outcome than failing a workload.

Upstream change Enforced from What a DCS cluster has to do
containerd 1.x support ends 1.36 Every node image must carry containerd 2.0 or later. Kubernetes 1.35 was the last release to support containerd 1.x, and from 1.38 an old containerd fails outright against the newer kubelet. Scrape kubelet_cri_losing_support before the upgrade to find nodes that are still behind.
cgroup v1 support phased out 1.35 Every node must run cgroup v2. failCgroupV1 has defaulted to true since 1.35, so the kubelet refuses to initialize on a cgroup v1 node. The failCgroupV1: false override still exists in 1.37, and it is only a stopgap.
Static Pods can no longer reference Secrets or ConfigMaps 1.37 Nothing on the DCS side. Static Pods below covers why, and names the one place the product touches the concept.
kube-proxy ipvs mode deprecated 1.37 logs a warning Confirm the mode with kubectl -n kube-system get configmap kube-proxy -o jsonpath='{.data.config\.conf}'. The mode is expected to be off by default in 1.40 and removed in 1.43, and the recommended Linux mode is now nftables. Talos and k3s both default to iptables, so a cluster nobody switched by hand is unaffected.
SELinuxMount graduates to GA 1.37 Only clusters with SELinux enabled see any effect, and only for CSI drivers that set seLinuxMount: true. Two pods with different SELinux labels sharing one volume can now fail to start where recursive relabeling used to let them coexist. Set seLinuxChangePolicy: Recursive on such a pod to keep the old behaviour.
metrics.k8s.io graduates to v1 1.37 Nothing. The gateway resolves the group version from discovery and prefers v1, falling back to v1beta1 on a cluster that serves only the beta (#1376). A cluster that later retires v1beta1 is followed, with no false report of a missing metrics stack.

Kubernetes v1.37 is planned for 26 August 2026. The v1.37 sneak peek is the source for every 1.37 row above.

The Talos version gates the Kubernetes version

On the reference deployments the Kubernetes version is not chosen independently. Talos v1.10 supports Kubernetes 1.28 through 1.33, so a cluster on Talos v1.10.9 cannot reach 1.37 at all until Talos itself is upgraded. Sequence the Talos bump first and the Kubernetes bump second.

That ordering also disposes of the first two rows of the table on Talos. Talos v1.10 already ships containerd 2.0.5, and the same release dropped cgroup v1 outside container mode. A Talos node current enough to run 1.37 satisfies both prerequisites by construction.

A Talos bump also moves two documentation links. The upstream disaster-recovery guide is linked at a version-pinned URL from the DR Runbook and from Backup and Recovery, because the Siderolabs documentation site serves no unversioned path for that page. Both links state the version they carry, and both name the release the reference deployments run. Neither one tracks the current Talos release. Move them with the bump, and confirm the commands quoted under each still match the new version's procedure.

Static Pods are not how DCS runs anything

The 1.37 kubelet prohibits a static Pod from referencing a Secret or a ConfigMap through fields like configMapRef or secretRef, and it removes the PreventStaticPodAPIReferences gate that used to let an operator opt out. Cloud-Native DCS ships no static Pods and writes nothing into /etc/kubernetes/manifests on any node. Unit runtime pods look node-pinned because they use hostNetwork and a hostPath state volume, but the physical-operator creates them through the API server like any other Pod (internal/controller/physical/unit_pod.go).

The product touches the static-Pod concept in exactly one place. The controller-removal drain skips pods carrying the kubernetes.io/config.mirror annotation, because a mirror pod belongs to the kubelet and cannot be evicted (internal/controller/physical/node_drain.go). Mirror-pod semantics are unchanged in 1.37, so that path needs no work.

Where the rule can still bite a plant is on a hand-rolled edge node. If a site runs a lightweight component as a static Pod alongside the DCS workload (a log shipper, a vendor agent, a bootstrap helper), read that manifest before the node crosses 1.37. Any configMapRef or secretRef content has to move inline into the manifest or onto the node's disk.

Pre-flight Checks

Run this checklist before touching the cluster. Anything that fails is a stop-ship.

# 1. Cluster and chart versions.
kubectl version --short
helm list -n dcs-system

# 2. All operators healthy and at the desired replica count.
kubectl get deploy -n dcs-system
kubectl get pods -n dcs-system \
  -l app.kubernetes.io/part-of=cloud-native-dcs

# 3. No batches in Running or Holding on a unit you plan to roll.
#    dcs is site-scoped — repeat per site (list sites with `dcs get sites`).
dcs get batches -s <site>
dcs get units -s <site>

# 4. Historian WAL replication is current (if historian is enabled).
kubectl cnpg status <historian-cluster-name>

# 5. Fresh etcd snapshot, younger than 1 hour.
talosctl -n <control-plane-ip> etcd snapshot pre-upgrade-$(date +%s).snapshot   # Talos
sudo k3s etcd-snapshot save pre-upgrade-$(date +%s)                             # k3s

# 6. Certificates are all Ready.
kubectl get certificate -n dcs-system \
  -l app.kubernetes.io/part-of=cloud-native-dcs

If the historian is configured with historian.backup.enabled: true, also confirm the most recent S3 backup is within its RPO window with kubectl get backup -n dcs-system.

If the window also moves the cluster to a new Kubernetes minor version, clear Kubernetes Version Upgrades before anything below. A node that fails its own prerequisites never gets as far as the chart.

If the deployment defines action-level authorization policies (gateway.auth.roles object form: allow/deny lists, ADR 0024), review the release notes' new-action list against your deny lists before upgrading: a new release can add actions to the catalog, and a deny list written against a family does not automatically cover a newly added sibling action. After the upgrade, dcs auth policy shows what the running gateway enforces.

A release can also move an existing action to a different tier, which changes who holds it without changing any role definition. The OPC UA discovery family (discovery:browse) moved from read to engineer in ADR 0062. Tooling that drove those routes under a credential below engineer has to move up, or be granted the action through an ADR 0024 allow entry on the role it holds. The same release added dcs-viewer to the shipped table, which is read alone and is the grant for an identity that only observes. A deployment supplying its own roles file does not receive it, since a non-empty gateway.auth.roles replaces the shipped table outright.

7. Dry-run the chart upgrade to catch immutable-field changes

The commands below use a local chart checkout. Upgrading against the published chart instead (oci://ghcr.io/cloud-native-dcs/charts/cloud-native-dcs) pulls from a private registry, so helm registry login ghcr.io has to have been run on the machine driving the upgrade. Credentials cached from an earlier release can be expired without any sign until the pull fails.

helm upgrade dcs deploy/helm/cloud-native-dcs \
  --namespace dcs-system --dry-run --reuse-values \
  --set global.image.tag=<target-version-tag>

The chart ships a render-time guard (templates/historian-database-immutability-check.yaml) that uses Helm's lookup function to compare the running CNPG Cluster's spec.postgresUID and spec.postgresGID against the values about to be applied. Both fields are immutable on an existing Cluster. Mismatching them otherwise trips the validating webhook, and Flux loops on RollbackFailed indefinitely. That failure mode was observed live during a 2026-04-22 demo-instance incident, which motivated this pre-flight check.

If --dry-run fails with an immutability error, the message includes the existing values and remediation. Honour it before retrying the live upgrade. The check is silent when nothing immutable changed.

The same dry-run also exercises the CNPG minimum-version guard (templates/historian-database-cnpg-version-check.yaml), which inspects the installed clusters.postgresql.cnpg.io CRD for .spec.podSecurityContext, a field the historian Cluster sets that only exists from CloudNativePG chart 0.27.0 (operator 1.28) onward. This matters most when upgrading a cluster whose CNPG was installed long ago against a loose version range: the DCS chart upgrade would otherwise fail mid-apply with .spec.podSecurityContext: field not declared in schema. Upgrade CNPG to >=0.27.0 first, then retry.

8. Validate PromQL changes against promtool

If the chart upgrade changes any rule under templates/prometheusrule.yaml, validate the rendered PromQL locally before the prometheus-operator's mutating admission webhook does. A syntax error there fails the entire Helm upgrade with Rules are not valid, and Flux gets stuck in Running 'upgrade' action until you manually unlock the release with flux suspend + helm rollback + flux resume.

helm template <release> deploy/helm/cloud-native-dcs/ \
  --set monitoring.prometheusRule.historianDisk.enabled=true \
  | python3 -c '
import sys, yaml
for d in yaml.safe_load_all(sys.stdin):
    if d and d.get("kind") == "PrometheusRule":
        print(yaml.safe_dump({"groups": d["spec"]["groups"]}))' \
  | docker run -i --entrypoint promtool \
      quay.io/prometheus/prometheus:v2.54.1 check rules /dev/stdin

helm lint and the chart's existing template tests do not catch PromQL syntax errors. Those are only surfaced by the in-cluster admission webhook. Add this check to your CI pipeline for any change under prometheusrule.yaml. See production-deployment.md § 4 "Validating PrometheusRule changes before deploy" for the recurring PromQL pitfalls (notably group_left() and the Helm-vs-Prometheus {{ }} overlap).

Upgrade Procedure

This sequence is the standard "rolling upgrade" path. Every step must complete successfully before moving to the next.

1. Drain or pause active batches

Rolling the operators and gateway doesn't interrupt running batches. State lives on CRDs, outside operator memory. Rolling a unit runtime pod does interrupt whatever is running on that unit. Two options:

  • Quiet window: wait for all batches to Complete, or Hold and Stop them via the HMI before starting.
  • Live upgrade: let the operators reconcile batches into Holding when their unit runtime goes NotReady. This is the default behaviour. The grace window before Hold is runtimeCrashGracePeriod in unit_controller.go. Expect each unit to lose ~30s of execution time.

Record the decision in the upgrade ticket before starting.

2. Apply the new CRDs

git fetch --tags origin
git checkout <target-version-tag>
make install        # runs kubectl apply -f config/crd/bases/
# or, directly:
kubectl apply -f config/crd/bases/

Watch for kubectl errors about incompatible schema changes -- those indicate a breaking CRD migration that the release notes should have flagged. If they didn't, stop and escalate.

3. Upgrade the chart

helm upgrade dcs deploy/helm/cloud-native-dcs \
  --namespace dcs-system \
  --reuse-values \
  --set global.image.tag=<target-version-tag> \
  --wait --timeout 10m

--wait blocks until every rollout reaches Ready. --reuse-values preserves any site-specific overrides you applied at install time. Pair it with --set for the small handful of values that need to change.

If you pin components by digest (Security Hardening § Supply Chain Verification), remember that a digest beats any tag: --reuse-values carries the old <component>.image.digest forward and that component will not move to <target-version-tag>. Update each pinned digest to the new release's digest in the same helm upgrade.

The MQTT broker is the one image that is not a DCS component. mqtt.image.tag names a Mosquitto release, and no longer the "2" series. An upgrade therefore moves it only when the chart's own default moves. That pin exists because the series tag changed the broker underneath a release nobody cut. A passwd file that made one Mosquitto accept the wrong credential made the next one terminate on load (#1579).

4. Verify each control-plane component rolled cleanly

kubectl rollout status deploy/dcs-cloud-native-dcs-physical-operator -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-procedural-operator -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-batch-operator -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-control-operator -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-gateway -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-historian -n dcs-system  # if enabled
kubectl rollout status deploy/dcs-cloud-native-dcs-mqtt -n dcs-system       # or sts in HA

Deployment names are <release>-cloud-native-dcs-<component> for a release whose name does not already contain the chart name. The commands on this page assume the release is named dcs, matching the helm upgrade dcs invocation above. The reference GitOps deployment pins releaseName: cloud-native-dcs, which collapses the prefix to cloud-native-dcs-<component>. kubectl get deploy -n dcs-system shows the rendered names for your install.

Operators use leader election, so the old leader steps down as soon as its replacement is Ready. The gateway is stateless.

5. Confirm the unit runtimes rolled, and finish the ones that could not

The physical-operator rolls a unit's runtime pod onto the new image as soon as it reconciles a unit that is free to be rolled, so most of this step is a check. Ask the product which units are still behind:

dcs --site <site> get units

Read the BUILD column. It is which build of the runtime is executing the control loop on that unit's edge node. It sits beside RUNTIME, which reports readiness. The two answer differently. A pod created by the previous release stays Running and Ready throughout. RUNTIME therefore reads Ready on every row, whether the upgrade finished or stopped at the control plane.

A site where every cell reads current is done. A cell reading blocked by batch is a unit allocated to a batch that is actively driving I/O, which is the one case the operator will not roll on its own: taking the runtime away mid-step can leave actuators in indeterminate states. The line under the table names those units.

Run it per site. dcs_unit_runtime_image_drift sums to zero when nothing anywhere is behind, and that is the one reading covering every site at once. The Unit detail page in the engineering UI carries the same verdict for one unit at a time.

Each of those has one ending, and it is to wait. When the batch reaches a resting phase, the next reconcile rolls the pod with no further gesture from anybody.

dcs runtime restart is not a second ending, and it is worth knowing why before you reach for it. It asks BatchPhase.IsResting about the same Batch the operator asks about. That means it refuses every unit the note lists, naming the batch and its phase. Once that batch does rest, the operator rolls the pod on its own within a reconcile. There is no state in which the command is the remedy, and the refusal you get instead is the product protecting the procedure.

Know what that costs before you plan the window. A unit allocated to a batch that will not rest for days keeps its runtime on the previous release for the whole of it. Ending the batch early is the only thing that changes the answer. Whether that is acceptable is a plant decision, and it is not one this procedure can make for you.

On the upgrade that first ships this behaviour, every runtime pod in every site predates it and is superseded by definition, so every free unit rolls at once on the new operator's first reconcile. Units are Idle during a planned window, which is what makes that acceptable. It does mean this upgrade recreates more pods at once than later ones will, and each recreate pulls the new image on its device node.

A chart upgrade landing on a plant that is mid-turnaround. The runtime image tag moves forward and every control-plane check goes green. The four idle units roll their own runtime pods onto the new image with no operator gesture. The one unit a clean-in-place batch is running on keeps the previous image, and dcs get units names it, says what will end the wait, and says the obvious remedy will refuse it. dcs runtime restart then does refuse it, for the same reason the operator declined. The closing beat is that wait. The batch finishes, the unit is released, and the next reconcile completes the roll with nobody typing anything. Both tags are the same build under two names on the capture stack, so no version moved.

6. Roll the io-probe pods

Nothing recreates a Running io-probe pod on a superseded image, and that is a decision. The probe reads every device it serves every 15 s, any telegram feeds that device's comm-loss watchdog, and the 45 s floor on spec.failSafe.timeout is three of those cadences. A recreate onto a new tag includes an image pull on a device node, so an unattended roll could put a real device into its fail-safe. The operator rolls these, deliberately, in a window of their choosing.

The product reports which probes are outstanding. Every IOModule carries an IOProbeImageDrift condition naming the pod that serves it, dcs_ioprobe_image_drift counts one series per probe, and dcs io list prints a PROBE column plus a line naming the distinct pods that need rolling:

dcs --site <site> io list

That line is deduped by pod, which the column cannot be. One probe serves every module on its Controller, so the number of restarts is the number of pods rather than the number of superseded cells. Run it per site. The IOModule detail page in the engineering UI carries the same verdict for one module at a time.

Check the declared fail-safes before rolling anything. An IOModule that declares spec.failSafe with a timeout can reach its fail-safe if the replacement pod is slow to start. The restart route refuses until you say you know, and it names the modules it is refusing about. Running it once without the acknowledgement is itself the check:

dcs --site <site> io probe restart <module> \
  --reason "roll the io-probe onto <version> after the chart upgrade"

Pre-pull the image on the device nodes, or pick a window where a device reaching its fail-safe is acceptable, then repeat with the acknowledgement:

dcs --site <site> io probe restart <module> \
  --reason "roll the io-probe onto <version> after the chart upgrade" \
  --acknowledge-fail-safe

One probe serves many modules, so one call rolls the pod for all of them. The command prints the images either side of the roll and names every module whose monitoring was interrupted. Repeat for one module per remaining probe until the condition above reports nothing.

The restart is recorded in the audit trail with the justification and the acknowledged exposure. kubectl delete pod -n <ns> -l dcs.io/component=io-probe still works and is neither audited nor refused, so keep it for a cluster whose gateway is down.

7. Verify the upgrade end to end

See Verification After Upgrade. Do not declare success until every item on that checklist is green.

Rollback Procedure

Rollback is only straightforward when the new version introduced no CRD schema changes. Check the release notes for Rollback-safe: yes before choosing this path.

Rollback path A -- rollback-safe upgrade

# 1. Helm rollback to the previous revision.
helm history dcs -n dcs-system
helm rollback dcs <previous-revision> -n dcs-system --wait --timeout 10m

# 2. The operator rolls each free unit's runtime pod back to the previous
#    image on its next reconcile. Confirm, and finish any unit a batch is
#    holding, exactly as in step 5 of the upgrade.

# 3. Roll the io-probe pods back. Nothing does this for you: one call per
#    probe, and the fail-safe acknowledgement applies in this direction too.
dcs --site <site> io probe restart <module> \
  --reason "roll the io-probe back to <previous version>" \
  --acknowledge-fail-safe

# 4. Run the verification checklist.

The rollback direction is not symmetric with the upgrade, and the asymmetry is worth knowing before you start. RuntimeImageDrift is a statement about disagreement between two images, and it says nothing about which of them is newer. The operator will roll a pod backwards onto the older image just as readily as forwards, because the image the chart now names is the image it will create a pod with. What that buys is a rollback that finishes itself for every free unit. What it costs is that a unit you deliberately left on the new runtime has nowhere to record that intent. Pin it with spec.runtimeImage on the Unit if you need one held.

IOProbeImageDrift is direction-blind for the same reason, and the probe half never finishes itself. Every probe still needs its own dcs io probe restart, and the condition reports the disagreement until one arrives.

helm rollback does not revert CRD changes applied with kubectl apply in step 2 of the upgrade. For rollback-safe upgrades this is fine. The older operators ignore new optional fields. For rollback-unsafe upgrades, use Path B.

Rollback path B -- CRD schema change or data-format migration

Once a new version has written CRs in a schema the old operators cannot read, helm rollback is not enough. You must restore etcd to the pre-upgrade snapshot:

  1. Stop the DCS workloads to prevent split-brain writes:
    kubectl scale deploy -n dcs-system \
      dcs-cloud-native-dcs-gateway dcs-cloud-native-dcs-physical-operator \
      dcs-cloud-native-dcs-procedural-operator dcs-cloud-native-dcs-batch-operator \
      dcs-cloud-native-dcs-control-operator --replicas=0
    
  2. Follow the etcd restore procedure in the DR Runbook for your platform: talosctl bootstrap --recover-from=pre-upgrade-<ts>.snapshot on Talos, or k3s server --cluster-reset --cluster-reset-restore-path=/var/lib/rancher/k3s/server/db/snapshots/pre-upgrade-<ts> on k3s.
  3. Re-apply the old CRDs from the version tag you are rolling back to:
    git checkout <previous-version-tag>
    kubectl apply -f config/crd/bases/
    
  4. helm rollback dcs <previous-revision> -n dcs-system.
  5. Run the runtime-pod roll from step 5 of the upgrade procedure.
  6. Run the verification checklist.

Data-format migrations (historian schema bumps, audit archive format changes) are one-way. If the upgrade ran the migration, a rollback loses any records written after the migration completed. The release notes for each release must call out which steps are one-way.

Downtime Expectations

Component Interruption during rolling upgrade
Gateway ~5-10s per replica as each pod goes NotReady and the service endpoints update. Active WebSocket subscriptions reconnect automatically. With gateway.replicas >= 2 and a PDB, user impact is effectively zero.
Operators None visible. State lives on CRDs; the new leader resumes reconciling where the old one left off.
MQTT broker ~5s reconnect on the clients. Telemetry queued in the runtime's store-and-forward buffer (queue.jsonl) replays automatically.
Historian None if CNPG-managed (no schema change). Historian pod restart briefly pauses ingest; MQTT QoS 1 redelivery backfills.
Unit runtime pods ~10-30s per pod from delete to new pod Ready. During this window the unit is NotReady; if a batch is Running on the unit, the grace window absorbs the gap. If the gap exceeds the grace window, the batch transitions to Hold and can be resumed.
Webhook (audit immutability) ~5s. During the restart window, AuditRecord create calls that race the webhook may be rejected; the gateway retries.
OMF egress Off by default, so most upgrades skip it. Where it is on, the pod is Recreate with one replica, so the gap is the pod restart. Its buffer is in memory: whatever it is holding for an unreachable endpoint at that moment is lost, and dcs_omf_egress_dropped_total records it. Upgrade while the endpoint is up. An outage it is riding out is the wrong window.

One-time restart on the upgrade that adds the OMF egress. The broker's auth Secret gains a fifth account, dcs-omf-egress, whether or not the component is enabled. The gateway, the four operators, and the historian each carry a checksum annotation over that Secret. Each therefore rolls exactly once, on that upgrade. It reads as an unexplained restart of components the release notes did not mention, and this is the reason.

Two things do not restart with them. The broker keeps running: it rehashes and SIGHUPs through its passwd-reloader sidecar, the same mechanism that makes password rotation graceful. Unit-runtime pods are recreated on a change to the runtime mTLS Secret and nothing else, and their own password did not change.

Verification After Upgrade

Every item must be green before closing the upgrade ticket.

  • [ ] dcs health returns all components healthy.
  • [ ] helm status dcs -n dcs-system shows the new chart version.
  • [ ] kubectl get deploy -n dcs-system shows the expected replica count for every DCS deployment.
  • [ ] Gateway is reachable and authenticates an OIDC user end to end.
  • [ ] No unit in any site is still on a superseded runtime image. This is the one item that covers every runtime pod in every site, where every other item here samples. Run dcs --site <site> get units per site and read the BUILD column: a finished site reads current on every row and prints no line under the table. dcs_unit_runtime_image_drift sums to zero when nothing anywhere is behind, and that is the one reading covering every site at once.
  • [ ] No IOModule in any site still reports a superseded probe. Nothing rolls a probe on its own, so this item is the whole of step 6's result. Run dcs --site <site> io list per site and read the PROBE column: a clean site prints no restart line under the table. dcs_ioprobe_image_drift sums to zero when nothing anywhere is superseded, and that is the one reading covering every site at once.
  • [ ] A canary batch runs from Create → Running → Complete against a simulated unit.
  • [ ] AuditRecords for the upgrade window are present and their e-signatures verify. See 21 CFR Part 11 traceability.
  • [ ] No x509, TLS handshake, or 401/403 spikes in gateway or operator logs in the 15 minutes after the rollout.
  • [ ] cert-manager is still healthy and all Certificate resources are Ready.
  • [ ] kubectl get prometheusrule -n dcs-system shows no firing alerts attributable to DCS.
  • Deploy Your Own -- initial install procedure
  • Backup and Recovery -- take backups before upgrading
  • DR Runbook -- recovery path when an upgrade corrupts state beyond what a rollback can fix
  • High Availability -- component failure modes during rolling upgrade
  • Rotation Runbook -- credential rotation is a frequent trigger for upgrade-time surprises
  • Historian Disk-Pressure Runbook -- upgrades that add new hypertables or change TimescaleDB versions can shift disk usage. Enable historian.prune on non-prod clusters to bound growth. Upgrades that touch historian.prune.priorityClassName or historian.prune.tolerateDiskPressure re-render the prune Pod template, which the next CronJob run picks up with no in-flight Pod restarts needed (issue #256).