Upgrade and Rollback¶
Status: Procedures below match the current Helm chart and the physical-operator's runtime-pod behaviour. The version compatibility matrix is a placeholder -- Cloud-Native DCS does not yet have stable release tags to map against, so that table lights up once v0.x releases exist.
Procedure for upgrading Cloud-Native DCS in place on an existing Kubernetes cluster and rolling back when an upgrade fails verification. Platform-specific steps (etcd snapshots) show both Talos (the reference deployments) and k3s forms.
Before You Start¶
- Read the release notes for the target version. Every release lists breaking changes, CRD schema migrations, and expected downtime.
- Confirm a current backup exists -- see Backup and Recovery. At minimum:
- etcd snapshot from the last hour:
talosctl -n <node> etcd snapshot pre-upgrade.snapshot(Talos) ork3s etcd-snapshot save(k3s) - CRD export:
dcs backup crds -o /backups/pre-upgrade-$(date +%F).yaml - HMAC signing key Secret:
kubectl get secret dcs-signing-key -n kube-system -o yaml > /backups/dcs-signing-key.yaml - Verify the current deployment is healthy. Upgrading on top of a
degraded system multiplies the failure modes you have to debug:
dcs health kubectl get pods -n dcs-system kubectl get certificate -n dcs-system -l app.kubernetes.io/part-of=cloud-native-dcs - Schedule the window. Even rolling upgrades briefly interrupt per-unit runtime pods (see Downtime Expectations). Plan for a quiet period when no batch is in Running or Holding on a unit you intend to roll.
Architecture Facts That Matter for Upgrades¶
Two chart behaviours tend to surprise first-time upgraders:
- Helm installs CRDs but never upgrades them. The chart ships the
CRDs in its
crds/directory, so a freshhelm installcreates them. Helm deliberately skips thecrds/directory onhelm upgrade, andhelm rollbackwill not undo a CRD change either. On every upgrade you must apply the new CRDs yourself withkubectl apply(ormake install, fromconfig/crd/bases/) before runninghelm upgrade. - Unit runtime pods and io-probe pods are reconciled by the operators
alone. Helm never touches them. They run as bare Pods in
site-<name>namespaces, so the chart upgrade rolls every Deployment and leaves these two carrying whatever image the operator that created them chose.
The two halves behave differently, and the difference is deliberate.
Unit runtime pods roll themselves where the plant allows it (#1762).
The physical-operator compares the running image against the one it would
create the pod with today. A unit is free when it has no allocated batch,
or when its batch is in a resting phase. For a free unit the operator
recreates the pod on the current image during its next reconcile, and
writes an AuditRecord naming both images. Where a batch is actively
driving I/O, it changes nothing and reports instead.
status.conditions[type=RuntimeImageDrift] goes True with reason
BlockedByBatch, naming the batch, and dcs_unit_runtime_image_drift
reads 1. Pulling the runtime out from under a Phase mid-step can leave
actuators in indeterminate states, which is the same reason
dcs runtime restart refuses it.
The runtime roll is therefore a step you verify. What remains for you is the units a batch was running on, and step 5 below is how you find those and account for them.
io-probe pods do not roll themselves. Nothing recreates a Running
probe pod on a superseded image, and that stays true on purpose: its read
cadence feeds every device comm-loss watchdog a module declares through
spec.failSafe, so an unattended roll is a plant action. The operator
reports the drift on every affected IOModule as an IOProbeImageDrift
condition, and dcs io probe restart is the audited remedy. Rolling
probes is step 6, and it is a deliberate gesture.
One-time migration: per-site runtime Certificate (v0.x, issue #213)¶
Upgrading from a chart version that predates #213 switches the CA issuer
from namespace-scoped Issuer to cluster-scoped ClusterIssuer and moves
per-site cloud-native-dcs-runtime-mtls Secret provisioning from a
gateway-managed copy to a cert-manager Certificate owned by the Site
reconciler. After helm upgrade:
# 1. Verify cert-manager is configured with --cluster-resource-namespace
# pointing at the DCS release namespace (usually dcs-system). Without
# this, cert-manager cannot find the CA Secret and ClusterIssuer stays
# NotReady.
kubectl get clusterissuer cloud-native-dcs-mtls-ca
# 2. Delete manually-copied runtime Secrets in each site namespace so the
# Site reconciler can recreate the Certificate → Secret chain cleanly:
for NS in $(kubectl get ns -l dcs.io/site -o name); do
kubectl delete secret -n "${NS#namespace/}" cloud-native-dcs-runtime-mtls --ignore-not-found
done
# 3. Trigger a Site reconcile so the reconciler provisions a Certificate.
# Any metadata change works — bump an annotation:
# kubectl annotate site <name> \
# dcs.io/reconcile-requested-at="$(date -u +%FT%TZ)" --overwrite
# cert-manager then issues a fresh Secret within seconds.
# 4. Bounce runtime pods once so they mount the fresh cert content:
kubectl get pod -A -l dcs.io/component=unit-runtime -o name | \
xargs -I {} kubectl delete {} -n <namespace>
Set mtls.certManager.clusterIssuer=false and mtls.certManager.perSiteRuntimeCert=false
if your cert-manager install can't be reconfigured. The pre-#213 behaviour
is preserved, but you keep the manual-copy drift risk.
Archive-integrity scheduler enabled by default (issue #216)¶
After upgrading to a chart that includes the Phase B scheduler, the
gateway runs dcs audit verify --archived on gateway.archiveIntegrity.interval
(default 6h) and writes one AuditRecord per run. Set
gateway.archiveIntegrity.enabled: false or env
GATEWAY_ARCHIVE_INTEGRITY_ENABLED=false to keep the old quiet
behaviour on downgrade-friendly test clusters. No schema change is
required. The scheduler reads the same audit_archive_manifest rows
written by the archiver in Phase A.
Historian disk-watchdog removed (issue #243)¶
The chart-shipped historian-disk-watchdog CronJob (the interim coverage
for #237/#238) has been removed. The two alerts it implemented are now
opt-in PrometheusRule resources gated on
monitoring.prometheusRule.historianDisk.enabled (default false). On
upgrade:
- Any value set under
historian.alerts.watchdog.*becomes a no-op. Helm silently ignores unknown keys, so the upgrade itself does not fail. The CronJob is removed from the cluster as part of the rollout. - If you rely on the watchdog's webhook, install
kube-prometheus-stack(or any Prometheus stack) and turn the new rules on:The reference deployment's Flux overlay (monitoring: prometheusRule: historianDisk: enabled: truecndcs-deploy-demo/flux/clusters/demo/) is the worked example. - The Slack/Discord/PagerDuty webhook URL is now consumed by Alertmanager,
not the CronJob. Move the
demo-alert-webhookSecret fromdcs-systeminto themonitoringnamespace, or recreate it there.
Production deployments should keep historianDisk.enabled: false and rely
on TimescaleDB native retention plus a CSI-backed PVC instead. See
production-deployment.md § 1, § 4.
Optional S3 Object-Lock mirror (issue #216, Phase C)¶
The audit-archiver chart now accepts
historian.audit.archival.immutable.* values. It is off by default.
Existing deployments keep the current PostgreSQL-only archive behaviour
until an operator explicitly points the archiver at a pre-provisioned
Object-Lock bucket plus a Secret with S3 credentials. Enabling mid-
cluster is safe. No migration is required, because PG remains the
primary query path and the mirror only covers batches written after the
flip. Bucket Object Lock cannot be added to an existing bucket
retroactively, so new buckets must be created with
--object-lock-enabled-for-bucket (AWS) or mc mb --with-lock
(minio). See backup-recovery.md
for the full checklist.
The mirror now needs its egress destination written down (issue #1514).
A deployment that already runs the mirror with networkPolicies.enabled
will find the upgrade refused at render time until
historian.audit.archival.immutable.egress.destinationCIDRs (or
destinationPodLabels, for a bucket served inside the cluster) is set. The
refusal names the value and an example. It is not a new requirement so much
as an old one that was never expressible: the archiver's NetworkPolicy has
always declared policyTypes: [Egress] and has never carried a rule for the
mirror endpoint, so on any cluster whose CNI enforces NetworkPolicy the
mirror upload was already being denied. The PostgreSQL archive kept
working, and only the 21 CFR Part 11 §11.10(c) restore copy went missing.
Nothing caught it because every stack the mirror had run on used a CNI that
ignores NetworkPolicy entirely.
Two things to check on the way through, and they are checked in different
places. Confirm the archiver's recent runs succeeded (kubectl get jobs -n
dcs-system over the CronJob's history), because a run that had a batch to
archive would have failed on the denied upload and left those AuditRecord
CRs in etcd, waiting for a run that can complete. Then list the bucket
itself. A period during which the archiver was denied is a period with no
immutable copy, and no later run rewrites it: a batch is mirrored once, by
the run that archives it. dcs audit verify --archived will not show that
hole, because it verifies the manifest chain in PostgreSQL and never reads
the bucket.
The historian database backup now needs its egress destination written down
too (issue #1516). A deployment that runs historian.backup.enabled with
networkPolicies.enabled will find the upgrade refused at render time
until historian.backup.s3.egress.destinationCIDRs (or
destinationPodLabels, for a store served inside the cluster) is set. It is
the same class of gap as the mirror above, in the policy one file over: the
CNPG pods' policy declares policyTypes: [Ingress, Egress] and its backup
rule named no destination at all, allowing every address in the world on
TCP/443, while the port the chart's own endpointURL example uses is 9000.
So a store on any port but 443 was already being denied. The port now
comes off endpointURL, and the hard-coded 443 is gone.
Check the store for a recent base backup on the way through, because a denial
here does not look like one. PostgreSQL does not drop a WAL segment it has
not archived, so pg_wal grows on the data volume until the volume fills.
That fill is what the 2026-04-12 demo incident was, one cause upstream.
IOModule session credentials move to a Secret (issue #1912)¶
An OPC UA IOModule authored before this release carried username and
password in spec.options. Those keys still open the session after the
upgrade. Nothing stops. Each such module reports
status.conditions[CredentialsSecured] as False/PlaintextInSpec, dcs io
get prints Credentials: PLAINTEXT in spec.options, and the gateway stops
serving the password on any read (ADR 0087).
Migrate a module by writing its credentials into a Secret in the site
namespace and replacing the option keys with a spec.security block that
names it (see Opening a secured session). The two are refused together
at admission, so the migration is one write. The unit-runtime pod serving
that module is recreated to mount the Secret, which is a runtime restart for
that unit. The io-probe is not recreated. The module reports
ProbeNotProjected until dcs io probe restart is run, and that is the
same read-gap decision step 6 below describes.
Once every module is on a Secret, set ioSecurity.refusePlaintextCredentials:
true in the chart values. From then on the driver refuses a module carrying
the deprecated keys in both pods, which is the declared boundary the
deployment asked for.
The IOModule CRD published with chart 0.7.3 does not install on a Kubernetes
1.31 apiserver: one of its validation rules spelled credentialsRef.namespace
where CEL needs __namespace__, and the apiserver refuses the whole CRD with
undefined field 'namespace'. A 1.33 or newer apiserver admits it. On an
older cluster, take the next patch release, whose CRD carries the escaped
spelling. Nothing else in 0.7.3 is affected, and a cluster that already holds
the CRD keeps it.
The pod reaper CronJob moves to Replace (issue #1937)¶
Charts before 0.7.6 ran the terminated-pod reaper under
concurrencyPolicy: Forbid. A reaper Job whose pod the scheduler refused
stayed active with nothing to end it, and Forbid then suppressed every later
run. Chart 0.6.3 added activeDeadlineSeconds to the job template, but that
field lives on the Job, and a Job is the CronJob's child. An upgrade rewrites
the CronJob and never touches a Job that already exists, so a cluster that was
already parked stayed parked through every upgrade since. On the reference
bench one Job created on 2026-08-25 survived five chart upgrades and nine days.
From 0.7.6 the CronJob runs under Replace. At the first :23 after the
upgrade the controller deletes any Job still active from the old chart and
creates the new run. There is nothing to run by hand, and no reaper run is
lost that the old policy would have kept. Confirm it afterwards:
kubectl -n dcs-system get cronjob -l app.kubernetes.io/component=pod-reaper
ACTIVE reads 0 between runs and LAST SCHEDULE is under an hour old. The
System Health page carries a Pod Reaper card from this release. It reads
degraded whenever the reaper has not succeeded in two schedules, so the
parked state is visible without reading the namespace by hand.
The archiver and prune CronJobs follow, and the prune alert reads silence (issue #1942)¶
The audit archiver and the historian prune CronJobs move to Replace in the
same release, and each Job template gains a deadline sized just under its own
schedule: 55 minutes for the hourly archiver
(historian.audit.archival.activeDeadlineSeconds) and five hours for the
six-hourly prune (historian.prune.activeDeadlineSeconds). A run parked on an
older chart is freed at the first tick after the upgrade, the same way as the
reaper. An archiver run killed at its deadline resumes on the next tick with
nothing lost, because every batch is durable before its records leave etcd. A
record the killed run had attested but not yet deleted is deleted by the next
run. Charts before 0.7.6 skipped such a record forever.
The Historian Prune card joins the System Health page. It reads degraded
when the prune has not succeeded in two of its six-hour schedules.
DCSHistorianPruneJobFailing no longer requires a failed Job. Charts before
0.7.6 keyed it on kube_job_status_failed. A Job whose pod never scheduled
never fails, so the rule could not see the parked shape at all. It fires on
silence now: no success in pruneStaleHours measured off the CronJob's own
last success, its creation when it has never succeeded, or any labelled Job's
completion. A suspended CronJob is exempt. The rule reads three series from
the kube-state-metrics cronjob collector that the old one did not,
kube_cronjob_status_last_successful_time, kube_cronjob_created and
kube_cronjob_spec_suspend. A kube-state-metrics too old to export them
leaves the rule reading the Job completions alone. A manual prune still moves
those.
Version Compatibility Matrix¶
Cloud-Native DCS is pre-1.0 and does not yet have release tags to map against. This table will light up as v0.x releases ship.
| From | To | Forward-compatible | Rollback-safe | Notes |
|---|---|---|---|---|
| pre-release | pre-release | n/a | n/a | No stable release yet |
When releases begin:
- Forward-compatible means the new operators tolerate the old CRD
schemas for one minor version, so a rolling upgrade is safe.
- Rollback-safe means the old operators can still read any CR that
was written by the new version. A rollback-unsafe upgrade typically
requires an etcd snapshot restore. A helm rollback cannot do it.
Kubernetes Version Upgrades¶
Upgrading the cluster underneath Cloud-Native DCS is a separate exercise from upgrading the chart, and every constraint on it comes from upstream. Cloud-Native DCS builds against the Kubernetes 1.36 API libraries and runs its envtest suites on the 1.36 control-plane binaries. The reference deployments run Talos Linux v1.10.9 with Kubernetes v1.33.6, and the documented floor for a customer cluster is Kubernetes 1.27.
The wave of upstream removals landing between Kubernetes 1.35 and 1.38 is about node prerequisites. None of it lands on the DCS control plane. Read the table before scheduling a cluster upgrade. Two of these rows fail the kubelet on a node, which is a worse outcome than failing a workload.
| Upstream change | Enforced from | What a DCS cluster has to do |
|---|---|---|
| containerd 1.x support ends | 1.36 | Every node image must carry containerd 2.0 or later. Kubernetes 1.35 was the last release to support containerd 1.x, and from 1.38 an old containerd fails outright against the newer kubelet. Scrape kubelet_cri_losing_support before the upgrade to find nodes that are still behind. |
| cgroup v1 support phased out | 1.35 | Every node must run cgroup v2. failCgroupV1 has defaulted to true since 1.35, so the kubelet refuses to initialize on a cgroup v1 node. The failCgroupV1: false override still exists in 1.37, and it is only a stopgap. |
| Static Pods can no longer reference Secrets or ConfigMaps | 1.37 | Nothing on the DCS side. Static Pods below covers why, and names the one place the product touches the concept. |
kube-proxy ipvs mode deprecated |
1.37 logs a warning | Confirm the mode with kubectl -n kube-system get configmap kube-proxy -o jsonpath='{.data.config\.conf}'. The mode is expected to be off by default in 1.40 and removed in 1.43, and the recommended Linux mode is now nftables. Talos and k3s both default to iptables, so a cluster nobody switched by hand is unaffected. |
SELinuxMount graduates to GA |
1.37 | Only clusters with SELinux enabled see any effect, and only for CSI drivers that set seLinuxMount: true. Two pods with different SELinux labels sharing one volume can now fail to start where recursive relabeling used to let them coexist. Set seLinuxChangePolicy: Recursive on such a pod to keep the old behaviour. |
metrics.k8s.io graduates to v1 |
1.37 | Nothing. The gateway resolves the group version from discovery and prefers v1, falling back to v1beta1 on a cluster that serves only the beta (#1376). A cluster that later retires v1beta1 is followed, with no false report of a missing metrics stack. |
Kubernetes v1.37 is planned for 26 August 2026. The v1.37 sneak peek is the source for every 1.37 row above.
The Talos version gates the Kubernetes version¶
On the reference deployments the Kubernetes version is not chosen independently. Talos v1.10 supports Kubernetes 1.28 through 1.33, so a cluster on Talos v1.10.9 cannot reach 1.37 at all until Talos itself is upgraded. Sequence the Talos bump first and the Kubernetes bump second.
That ordering also disposes of the first two rows of the table on Talos. Talos v1.10 already ships containerd 2.0.5, and the same release dropped cgroup v1 outside container mode. A Talos node current enough to run 1.37 satisfies both prerequisites by construction.
A Talos bump also moves two documentation links. The upstream disaster-recovery guide is linked at a version-pinned URL from the DR Runbook and from Backup and Recovery, because the Siderolabs documentation site serves no unversioned path for that page. Both links state the version they carry, and both name the release the reference deployments run. Neither one tracks the current Talos release. Move them with the bump, and confirm the commands quoted under each still match the new version's procedure.
Static Pods are not how DCS runs anything¶
The 1.37 kubelet prohibits a static Pod from referencing a Secret or a
ConfigMap through fields like configMapRef or secretRef, and it removes
the PreventStaticPodAPIReferences gate that used to let an operator opt
out. Cloud-Native DCS ships no static Pods and writes nothing into
/etc/kubernetes/manifests on any node. Unit runtime pods look node-pinned
because they use hostNetwork and a hostPath state volume, but the
physical-operator creates them through the API server like any other Pod
(internal/controller/physical/unit_pod.go).
The product touches the static-Pod concept in exactly one place. The
controller-removal drain skips pods carrying the
kubernetes.io/config.mirror annotation, because a mirror pod belongs to
the kubelet and cannot be evicted
(internal/controller/physical/node_drain.go). Mirror-pod semantics are
unchanged in 1.37, so that path needs no work.
Where the rule can still bite a plant is on a hand-rolled edge node. If a
site runs a lightweight component as a static Pod alongside the DCS
workload (a log shipper, a vendor agent, a bootstrap helper), read that
manifest before the node crosses 1.37. Any configMapRef or secretRef
content has to move inline into the manifest or onto the node's disk.
Pre-flight Checks¶
Run this checklist before touching the cluster. Anything that fails is a stop-ship.
# 1. Cluster and chart versions.
kubectl version --short
helm list -n dcs-system
# 2. All operators healthy and at the desired replica count.
kubectl get deploy -n dcs-system
kubectl get pods -n dcs-system \
-l app.kubernetes.io/part-of=cloud-native-dcs
# 3. No batches in Running or Holding on a unit you plan to roll.
# dcs is site-scoped — repeat per site (list sites with `dcs get sites`).
dcs get batches -s <site>
dcs get units -s <site>
# 4. Historian WAL replication is current (if historian is enabled).
kubectl cnpg status <historian-cluster-name>
# 5. Fresh etcd snapshot, younger than 1 hour.
talosctl -n <control-plane-ip> etcd snapshot pre-upgrade-$(date +%s).snapshot # Talos
sudo k3s etcd-snapshot save pre-upgrade-$(date +%s) # k3s
# 6. Certificates are all Ready.
kubectl get certificate -n dcs-system \
-l app.kubernetes.io/part-of=cloud-native-dcs
If the historian is configured with historian.backup.enabled: true,
also confirm the most recent S3 backup is within its RPO window with
kubectl get backup -n dcs-system.
If the window also moves the cluster to a new Kubernetes minor version, clear Kubernetes Version Upgrades before anything below. A node that fails its own prerequisites never gets as far as the chart.
If the deployment defines action-level authorization policies
(gateway.auth.roles object form: allow/deny lists,
ADR 0024), review the
release notes' new-action list against your deny lists before
upgrading: a new release can add actions to the catalog, and a deny list
written against a family does not automatically cover a newly added
sibling action. After the upgrade, dcs auth policy shows what the
running gateway enforces.
A release can also move an existing action to a different tier, which
changes who holds it without changing any role definition. The OPC UA
discovery family (discovery:browse) moved from read to engineer in
ADR 0062. Tooling that
drove those routes under a credential below engineer has to move up, or
be granted the action through an ADR 0024 allow entry on the role it
holds. The same release added dcs-viewer to the shipped table, which is
read alone and is the grant for an identity that only observes. A
deployment supplying its own roles file does not receive it, since a
non-empty gateway.auth.roles replaces the shipped table outright.
7. Dry-run the chart upgrade to catch immutable-field changes¶
The commands below use a local chart checkout. Upgrading against the
published chart instead (oci://ghcr.io/cloud-native-dcs/charts/cloud-native-dcs)
pulls from a private registry, so helm registry login ghcr.io has to have
been run on the machine driving the upgrade. Credentials cached from an
earlier release can be expired without any sign until the pull fails.
helm upgrade dcs deploy/helm/cloud-native-dcs \
--namespace dcs-system --dry-run --reuse-values \
--set global.image.tag=<target-version-tag>
The chart ships a render-time guard
(templates/historian-database-immutability-check.yaml)
that uses Helm's lookup function to compare the running CNPG Cluster's
spec.postgresUID and spec.postgresGID against the values about to be
applied. Both fields are immutable on an existing Cluster. Mismatching
them otherwise trips the validating webhook, and Flux loops on
RollbackFailed indefinitely. That failure mode was observed live during
a 2026-04-22 demo-instance incident, which motivated this pre-flight
check.
If --dry-run fails with an immutability error, the message includes
the existing values and remediation. Honour it before retrying the live
upgrade. The check is silent when nothing immutable changed.
The same dry-run also exercises the CNPG minimum-version guard
(templates/historian-database-cnpg-version-check.yaml),
which inspects the installed clusters.postgresql.cnpg.io CRD for
.spec.podSecurityContext, a field the historian Cluster sets that
only exists from CloudNativePG chart 0.27.0 (operator 1.28) onward. This
matters most when upgrading a cluster whose CNPG was installed long ago
against a loose version range: the DCS chart upgrade would otherwise fail
mid-apply with .spec.podSecurityContext: field not declared in schema.
Upgrade CNPG to >=0.27.0 first, then retry.
8. Validate PromQL changes against promtool¶
If the chart upgrade changes any rule under
templates/prometheusrule.yaml, validate the rendered PromQL locally
before the prometheus-operator's mutating admission webhook does. A
syntax error there fails the entire Helm upgrade with Rules are not
valid, and Flux gets stuck in Running 'upgrade' action until you
manually unlock the release with flux suspend + helm rollback +
flux resume.
helm template <release> deploy/helm/cloud-native-dcs/ \
--set monitoring.prometheusRule.historianDisk.enabled=true \
| python3 -c '
import sys, yaml
for d in yaml.safe_load_all(sys.stdin):
if d and d.get("kind") == "PrometheusRule":
print(yaml.safe_dump({"groups": d["spec"]["groups"]}))' \
| docker run -i --entrypoint promtool \
quay.io/prometheus/prometheus:v2.54.1 check rules /dev/stdin
helm lint and the chart's existing template tests do not catch
PromQL syntax errors. Those are only surfaced by the in-cluster
admission webhook. Add this check to your CI pipeline for any change
under prometheusrule.yaml. See
production-deployment.md § 4 "Validating PrometheusRule changes
before deploy"
for the recurring PromQL pitfalls (notably group_left() and the
Helm-vs-Prometheus {{ }} overlap).
Upgrade Procedure¶
This sequence is the standard "rolling upgrade" path. Every step must complete successfully before moving to the next.
1. Drain or pause active batches¶
Rolling the operators and gateway doesn't interrupt running batches. State lives on CRDs, outside operator memory. Rolling a unit runtime pod does interrupt whatever is running on that unit. Two options:
- Quiet window: wait for all batches to Complete, or Hold and Stop them via the HMI before starting.
- Live upgrade: let the operators reconcile batches into Holding
when their unit runtime goes NotReady. This is the default behaviour.
The grace window before Hold is
runtimeCrashGracePeriodinunit_controller.go. Expect each unit to lose ~30s of execution time.
Record the decision in the upgrade ticket before starting.
2. Apply the new CRDs¶
git fetch --tags origin
git checkout <target-version-tag>
make install # runs kubectl apply -f config/crd/bases/
# or, directly:
kubectl apply -f config/crd/bases/
Watch for kubectl errors about incompatible schema changes -- those
indicate a breaking CRD migration that the release notes should have
flagged. If they didn't, stop and escalate.
3. Upgrade the chart¶
helm upgrade dcs deploy/helm/cloud-native-dcs \
--namespace dcs-system \
--reuse-values \
--set global.image.tag=<target-version-tag> \
--wait --timeout 10m
--wait blocks until every rollout reaches Ready. --reuse-values
preserves any site-specific overrides you applied at install time. Pair
it with --set for the small handful of values that need to change.
If you pin components by digest
(Security Hardening § Supply Chain Verification),
remember that a digest beats any tag: --reuse-values carries the old
<component>.image.digest forward and that component will not move to
<target-version-tag>. Update each pinned digest to the new release's
digest in the same helm upgrade.
The MQTT broker is the one image that is not a DCS component. mqtt.image.tag
names a Mosquitto release, and no longer the "2" series. An upgrade therefore
moves it only when the chart's own default moves. That pin exists because the
series
tag changed the broker underneath a release nobody cut. A passwd file that made
one Mosquitto accept the wrong credential made the next one terminate on load
(#1579).
4. Verify each control-plane component rolled cleanly¶
kubectl rollout status deploy/dcs-cloud-native-dcs-physical-operator -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-procedural-operator -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-batch-operator -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-control-operator -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-gateway -n dcs-system
kubectl rollout status deploy/dcs-cloud-native-dcs-historian -n dcs-system # if enabled
kubectl rollout status deploy/dcs-cloud-native-dcs-mqtt -n dcs-system # or sts in HA
Deployment names are <release>-cloud-native-dcs-<component> for a release
whose name does not already contain the chart name. The commands on this
page assume the release is named dcs, matching the helm upgrade dcs
invocation above. The reference GitOps deployment pins
releaseName: cloud-native-dcs, which collapses the prefix to
cloud-native-dcs-<component>. kubectl get deploy -n dcs-system shows
the rendered names for your install.
Operators use leader election, so the old leader steps down as soon as its replacement is Ready. The gateway is stateless.
5. Confirm the unit runtimes rolled, and finish the ones that could not¶
The physical-operator rolls a unit's runtime pod onto the new image as soon as it reconciles a unit that is free to be rolled, so most of this step is a check. Ask the product which units are still behind:
dcs --site <site> get units
Read the BUILD column. It is which build of the runtime is executing the
control loop on that unit's edge node. It sits beside RUNTIME, which reports
readiness. The two answer differently. A pod created by the previous
release stays Running and Ready throughout. RUNTIME therefore reads Ready
on every row, whether the upgrade finished or stopped at the control plane.
A site where every cell reads current is done. A cell reading
blocked by batch is a unit allocated to a batch that is actively driving
I/O, which is the one case the operator will not roll on its own: taking the
runtime away mid-step can leave actuators in indeterminate states. The line
under the table names those units.
Run it per site. dcs_unit_runtime_image_drift sums to zero when nothing
anywhere is behind, and that is the one reading covering every site at once.
The Unit detail page in the engineering UI carries the same verdict for one
unit at a time.
Each of those has one ending, and it is to wait. When the batch reaches a resting phase, the next reconcile rolls the pod with no further gesture from anybody.
dcs runtime restart is not a second ending, and it is worth knowing why
before you reach for it. It asks BatchPhase.IsResting about the same Batch
the operator asks about. That means it refuses every unit the note lists,
naming the batch and its phase. Once that batch does rest, the operator rolls
the pod on its own within a reconcile. There is no state in which the command
is the remedy, and the refusal you get instead is the product protecting the
procedure.
Know what that costs before you plan the window. A unit allocated to a batch that will not rest for days keeps its runtime on the previous release for the whole of it. Ending the batch early is the only thing that changes the answer. Whether that is acceptable is a plant decision, and it is not one this procedure can make for you.
On the upgrade that first ships this behaviour, every runtime pod in every site predates it and is superseded by definition, so every free unit rolls at once on the new operator's first reconcile. Units are Idle during a planned window, which is what makes that acceptable. It does mean this upgrade recreates more pods at once than later ones will, and each recreate pulls the new image on its device node.
dcs get units names it, says what will end
the wait, and says the obvious remedy will refuse it.
dcs runtime restart then does refuse it, for the same reason the
operator declined. The closing beat is that wait. The batch finishes, the unit
is released, and the next reconcile completes the roll with nobody typing
anything. Both tags are the same build under two names on the
capture stack, so no version moved.6. Roll the io-probe pods¶
Nothing recreates a Running io-probe pod on a superseded image, and that is a
decision. The probe reads every device it serves every 15 s,
any telegram feeds that device's comm-loss watchdog, and the 45 s floor on
spec.failSafe.timeout is three of those cadences. A recreate onto a new tag
includes an image pull on a device node, so an unattended roll could put a real
device into its fail-safe. The operator rolls these, deliberately, in a window
of their choosing.
The product reports which probes are outstanding. Every IOModule carries an
IOProbeImageDrift condition naming the pod that serves it,
dcs_ioprobe_image_drift counts one series per probe, and dcs io list prints
a PROBE column plus a line naming the distinct pods that need rolling:
dcs --site <site> io list
That line is deduped by pod, which the column cannot be. One probe serves every module on its Controller, so the number of restarts is the number of pods rather than the number of superseded cells. Run it per site. The IOModule detail page in the engineering UI carries the same verdict for one module at a time.
Check the declared fail-safes before rolling anything. An IOModule that
declares spec.failSafe with a timeout can reach its fail-safe if the
replacement pod is slow to start. The restart route refuses until you say you know, and it names the modules it is
refusing about. Running it once without the acknowledgement is itself the check:
dcs --site <site> io probe restart <module> \
--reason "roll the io-probe onto <version> after the chart upgrade"
Pre-pull the image on the device nodes, or pick a window where a device reaching its fail-safe is acceptable, then repeat with the acknowledgement:
dcs --site <site> io probe restart <module> \
--reason "roll the io-probe onto <version> after the chart upgrade" \
--acknowledge-fail-safe
One probe serves many modules, so one call rolls the pod for all of them. The command prints the images either side of the roll and names every module whose monitoring was interrupted. Repeat for one module per remaining probe until the condition above reports nothing.
The restart is recorded in the audit trail with the justification and the
acknowledged exposure. kubectl delete pod -n <ns> -l dcs.io/component=io-probe
still works and is neither audited nor refused, so keep it for a cluster whose
gateway is down.
7. Verify the upgrade end to end¶
See Verification After Upgrade. Do not declare success until every item on that checklist is green.
Rollback Procedure¶
Rollback is only straightforward when the new version introduced no CRD
schema changes. Check the release notes for Rollback-safe: yes before
choosing this path.
Rollback path A -- rollback-safe upgrade¶
# 1. Helm rollback to the previous revision.
helm history dcs -n dcs-system
helm rollback dcs <previous-revision> -n dcs-system --wait --timeout 10m
# 2. The operator rolls each free unit's runtime pod back to the previous
# image on its next reconcile. Confirm, and finish any unit a batch is
# holding, exactly as in step 5 of the upgrade.
# 3. Roll the io-probe pods back. Nothing does this for you: one call per
# probe, and the fail-safe acknowledgement applies in this direction too.
dcs --site <site> io probe restart <module> \
--reason "roll the io-probe back to <previous version>" \
--acknowledge-fail-safe
# 4. Run the verification checklist.
The rollback direction is not symmetric with the upgrade, and the asymmetry
is worth knowing before you start. RuntimeImageDrift is a statement about
disagreement between two images, and it says nothing about which of them is
newer. The operator will roll a pod backwards onto the older image just as
readily as forwards, because the image the chart now names is the image it
will create a pod with. What that buys is a rollback that finishes itself for
every free unit. What it costs is that a unit you deliberately left on the new
runtime has nowhere to record that intent. Pin it with spec.runtimeImage on
the Unit if you need one held.
IOProbeImageDrift is direction-blind for the same reason, and the probe half
never finishes itself. Every probe still needs its own dcs io probe restart,
and the condition reports the disagreement until one arrives.
helm rollback does not revert CRD changes applied with
kubectl apply in step 2 of the upgrade. For rollback-safe upgrades this
is fine. The older operators ignore new optional fields. For
rollback-unsafe upgrades, use Path B.
Rollback path B -- CRD schema change or data-format migration¶
Once a new version has written CRs in a schema the old operators cannot
read, helm rollback is not enough. You must restore etcd to the
pre-upgrade snapshot:
- Stop the DCS workloads to prevent split-brain writes:
kubectl scale deploy -n dcs-system \ dcs-cloud-native-dcs-gateway dcs-cloud-native-dcs-physical-operator \ dcs-cloud-native-dcs-procedural-operator dcs-cloud-native-dcs-batch-operator \ dcs-cloud-native-dcs-control-operator --replicas=0 - Follow the etcd restore procedure in the
DR Runbook for your platform:
talosctl bootstrap --recover-from=pre-upgrade-<ts>.snapshoton Talos, ork3s server --cluster-reset --cluster-reset-restore-path=/var/lib/rancher/k3s/server/db/snapshots/pre-upgrade-<ts>on k3s. - Re-apply the old CRDs from the version tag you are rolling back to:
git checkout <previous-version-tag> kubectl apply -f config/crd/bases/ helm rollback dcs <previous-revision> -n dcs-system.- Run the runtime-pod roll from step 5 of the upgrade procedure.
- Run the verification checklist.
Data-format migrations (historian schema bumps, audit archive format changes) are one-way. If the upgrade ran the migration, a rollback loses any records written after the migration completed. The release notes for each release must call out which steps are one-way.
Downtime Expectations¶
| Component | Interruption during rolling upgrade |
|---|---|
| Gateway | ~5-10s per replica as each pod goes NotReady and the service endpoints update. Active WebSocket subscriptions reconnect automatically. With gateway.replicas >= 2 and a PDB, user impact is effectively zero. |
| Operators | None visible. State lives on CRDs; the new leader resumes reconciling where the old one left off. |
| MQTT broker | ~5s reconnect on the clients. Telemetry queued in the runtime's store-and-forward buffer (queue.jsonl) replays automatically. |
| Historian | None if CNPG-managed (no schema change). Historian pod restart briefly pauses ingest; MQTT QoS 1 redelivery backfills. |
| Unit runtime pods | ~10-30s per pod from delete to new pod Ready. During this window the unit is NotReady; if a batch is Running on the unit, the grace window absorbs the gap. If the gap exceeds the grace window, the batch transitions to Hold and can be resumed. |
| Webhook (audit immutability) | ~5s. During the restart window, AuditRecord create calls that race the webhook may be rejected; the gateway retries. |
| OMF egress | Off by default, so most upgrades skip it. Where it is on, the pod is Recreate with one replica, so the gap is the pod restart. Its buffer is in memory: whatever it is holding for an unreachable endpoint at that moment is lost, and dcs_omf_egress_dropped_total records it. Upgrade while the endpoint is up. An outage it is riding out is the wrong window. |
One-time restart on the upgrade that adds the OMF egress. The broker's
auth Secret gains a fifth account, dcs-omf-egress, whether or not the
component is enabled. The gateway, the four operators, and the historian each
carry a checksum annotation over that Secret. Each therefore rolls exactly
once, on that upgrade. It reads as an unexplained restart of
components the release notes did not mention, and this is the reason.
Two things do not restart with them. The broker keeps running: it rehashes
and SIGHUPs through its passwd-reloader sidecar, the same mechanism
that makes password rotation graceful. Unit-runtime pods are recreated on a
change to the runtime mTLS Secret and nothing else, and their own password did
not change.
Verification After Upgrade¶
Every item must be green before closing the upgrade ticket.
- [ ]
dcs healthreturns all components healthy. - [ ]
helm status dcs -n dcs-systemshows the new chart version. - [ ]
kubectl get deploy -n dcs-systemshows the expected replica count for every DCS deployment. - [ ] Gateway is reachable and authenticates an OIDC user end to end.
- [ ] No unit in any site is still on a superseded runtime image. This is the
one item that covers every runtime pod in every site, where every other
item here samples. Run
dcs --site <site> get unitsper site and read theBUILDcolumn: a finished site readscurrenton every row and prints no line under the table.dcs_unit_runtime_image_driftsums to zero when nothing anywhere is behind, and that is the one reading covering every site at once. - [ ] No IOModule in any site still reports a superseded probe. Nothing rolls a
probe on its own, so this item is the whole of step 6's result. Run
dcs --site <site> io listper site and read thePROBEcolumn: a clean site prints no restart line under the table.dcs_ioprobe_image_driftsums to zero when nothing anywhere is superseded, and that is the one reading covering every site at once. - [ ] A canary batch runs from Create → Running → Complete against a simulated unit.
- [ ] AuditRecords for the upgrade window are present and their e-signatures verify. See 21 CFR Part 11 traceability.
- [ ] No
x509,TLS handshake, or401/403spikes in gateway or operator logs in the 15 minutes after the rollout. - [ ] cert-manager is still healthy and all
Certificateresources are Ready. - [ ]
kubectl get prometheusrule -n dcs-systemshows no firing alerts attributable to DCS.
Related Documentation¶
- Deploy Your Own -- initial install procedure
- Backup and Recovery -- take backups before upgrading
- DR Runbook -- recovery path when an upgrade corrupts state beyond what a rollback can fix
- High Availability -- component failure modes during rolling upgrade
- Rotation Runbook -- credential rotation is a frequent trigger for upgrade-time surprises
- Historian Disk-Pressure Runbook -- upgrades
that add new hypertables or change TimescaleDB versions can shift disk
usage. Enable
historian.pruneon non-prod clusters to bound growth. Upgrades that touchhistorian.prune.priorityClassNameorhistorian.prune.tolerateDiskPressurere-render the prune Pod template, which the next CronJob run picks up with no in-flight Pod restarts needed (issue #256).