# Helm Chart Upgrades — 2026-07-12 ## Goal Upgrade all Helm-managed charts on the RKE2 cluster after the v1.35.6 upgrade. 5 components needed updates; RKE2-bundled charts were already current. ## Version Matrix (Final) | Component | Old Version | New Version | Method | Status | |-----------|------------|-------------|--------|--------| | External Secrets | 2.2.0 | 2.7.0 | Helm direct | ✅ | | CNPG | 0.22.1 (app 1.24.1) | 0.29.0 (app 1.30.0) | ArgoCD GitOps | ✅ | | Velero | 8.1.0 (app 1.15.0) | 12.1.0 (app 1.18.1) | ArgoCD GitOps | ✅ | | Ceph CSI RBD | 3.10.1 | 3.17.0 | Helm direct | ✅ | | ArgoCD | 7.8.13 (app v2.14.7) | 10.1.3 (app v3.4.5) | Helm direct | ✅ | | Cilium | rke2-cilium-1.19.402 | — | RKE2-bundled | Already current | | Traefik | rke2-traefik-40.1.003 | — | RKE2-bundled | Already current | | CoreDNS | rke2-coredns-1.46.002 | — | RKE2-bundled | Already current | | Metrics Server | rke2-metrics-server-3.13.100 | — | RKE2-bundled | Already current | | Snapshot Ctrl | rke2-snapshot-controller-4.2.006 | — | RKE2-bundled | Already current | ## Upgrade Order (Safest → Riskiest) 1. External Secrets 2.2.0 → 2.7.0 (Minor, Helm direct, minimal values) 2. CNPG 0.22.1 → 0.29.0 (Major, ArgoCD GitOps, **required replica fix first**) 3. Velero 8.1.0 → 12.1.0 (Major, ArgoCD GitOps) 4. Ceph CSI RBD 3.10.1 → 3.17.0 (Major, Helm direct, **causes transient API outage**) 5. ArgoCD v2.14.7 → v3.4.5 (MAJOR v2→v3, Helm direct, **breaking change**) ## Detail: External Secrets (2.2.0 → 2.7.0) Straightforward `helm upgrade --reuse-values`. Values were minimal (`installCRDs: true`). No breaking changes. All ExternalSecrets continued syncing. Two pre-existing broken ExternalSecrets (argocd-repo-credentials, postgres-main-db) were already broken for 110+ days — unrelated to the upgrade. ## Detail: CNPG (0.22.1 → 0.29.0) — Replica Rebuild Required ### Pre-Existing Issue Found `postgres-main-2` replica was broken for 109 days: ``` requested WAL segment 000000020000000000000008 has already been removed WAL file not found in the recovery object store ``` Primary had recycled WAL segments the replica still needed. Cluster was running with 2/3 Ready (primary + 1 replica), backups operational. ### Replica Rebuild ```bash # Delete broken replica PVCs + pod kubectl delete pvc postgres-main-2 postgres-main-2-wal -n postgres kubectl delete pod postgres-main-2 -n postgres # CNPG automatically created postgres-main-4 as replacement # Wait for 3/3 Ready: kubectl get pods -n postgres -w ``` ### Operator Upgrade via GitOps ```bash # Changed targetRevision in ArgoCD Application manifest cd /root/iac-homelab sed -i 's/targetRevision: 0.22.1/targetRevision: 0.29.0/' \ clusters/main/operators/cloudnativepg.yaml git add -A && git commit -m "feat: upgrade CNPG 0.22.1 → 0.29.0" && git push ``` **Pitfall**: ArgoCD didn't auto-detect the chart version change despite `selfHeal: true`. Required hard refresh: ```bash kubectl annotate application cloudnativepg-operator -n argocd \ argocd.argoproj.io/refresh=hard --overwrite ``` After refresh, ArgoCD pulled new chart, created new operator pod with image v1.30.0. Postgres cluster briefly went 2/3 during reconciliation (operator triggered primary restart), then recovered to 3/3 Healthy. ## Detail: Velero (8.1.0 → 12.1.0) Same GitOps pattern as CNPG — changed `targetRevision` in ArgoCD Application manifest, pushed, hard-refreshed ArgoCD. All Velero components rolled out: server + 6 node-agents (one per K8s node). BSL (Backup Storage Location) remained Available throughout. Backup schedules (daily-full-cluster, weekly-full-cluster) continued working. ## Detail: Ceph CSI RBD (3.10.1 → 3.17.0) — API Server Disruption ### Upgrade ```bash helm upgrade ceph-csi-rbd cephcsi/ceph-csi-rbd \ --version 3.17.0 -n kube-system --reuse-values --timeout 120s ``` ### Transient API Server Outage The nodeplugin DaemonSet rolling update disrupted kubelet on CP nodes. All 3 API servers became unreachable: ``` The connection to the server 10.0.30.51:6443 was refused ``` Recovery took ~5 minutes. Monitored with: ```bash curl -sk --connect-timeout 5 https://10.0.30.51:6443/healthz # 401 Unauthorized = server is UP (expected without credentials) # Connection refused = server is DOWN ``` After recovery, provisioner pods (2 replicas) were 7/7 Running with image `quay.io/cephcsi/cephcsi:v3.17.0`. One nodeplugin pod stayed in `ContainerCreating` for several minutes (image pull delay on CP-01). All 7 PVCs remained bound throughout — no data impact. ### Root Cause The Ceph CSI nodeplugin DaemonSet runs on EVERY node including CP nodes. Rolling the DaemonSet causes the CSI driver socket to be briefly unavailable, which can block kubelet operations on CP nodes, temporarily preventing the API server from serving requests. ## Detail: ArgoCD v2.14.7 → v3.4.5 (MAJOR Upgrade) ### Safe Path: Two-Step Upgrade Step 1: Patch to latest v2.x (low risk): ```bash helm upgrade argocd argo/argo-cd --version 7.9.1 -n argocd --reuse-values # v2.14.7 → v2.14.11 ``` Step 2: Major v2→v3 upgrade with explicit values: ```bash # v3 chart requires redis.networkPolicy.create field (breaking change) # --reuse-values FAILS with: # nil pointer evaluating interface {}.create # Must provide explicit values file with redis.networkPolicy.create: false helm upgrade argocd argo/argo-cd --version 10.1.3 -n argocd \ -f /tmp/argocd-v3-values.yaml --timeout 300s ``` ### v3 Values File Key addition vs v2 values: ```yaml redis: networkPolicy: create: false ``` All other values carried over unchanged (server.insecure, ingress, replicas, affinity). ### Post-Upgrade Verification - All ArgoCD pods Running (server ×2, repo-server ×2, application-controller, applicationset-controller, dex-server, notifications-controller, redis) - Image: `quay.io/argoproj/argocd:v3.4.5` - All applications Synced (except `backups` — pre-existing ExternalSecret OutOfSync, see §13.7) ## Detail: ArgoCD `backups` App OutOfSync (Task 4) ### Root Cause Single resource OutOfSync: `ExternalSecret/velero-s3-credentials` in `velero` namespace. ESO adds default fields to the live spec that aren't in the Git YAML: - `conversionStrategy: "Default"` - `decodingStrategy: "None"` - `metadataPolicy: "None"` - `deletionPolicy: "Retain"` - `engineVersion: "v2"` - `mergePolicy: "Replace"` ArgoCD's `ignoreDifferences` only covers `/status`, not `/spec`. ### Impact Cosmetic — ExternalSecret functions correctly (`secret synced`, Velero backups operational). No data or functionality impact. ### Fix (Not Yet Applied — Read-Only Investigation) Add ESO default fields to Git YAML at `clusters/main/backups/velero/external-secret.yaml` so spec matches exactly. Alternative: expand `ignoreDifferences` to cover specific JSON pointers (less clean). ## GitOps Trail Commits pushed to `dominik/iac-homelab` on Gitea: - `36eb5c4` — `feat(helm): upgrade all managed charts` (values files + ArgoCD Application targetRevision changes) - `acd3231` — `chore: remove .kube cache from repo, add to .gitignore` (cleanup of accidentally committed `.kube/cache/` directory) ## Lessons Learned 1. **Always check CNPG cluster health before upgrading the operator** — a broken replica (WAL gap) should be rebuilt first (delete PVC + pod, CNPG auto-creates fresh replica from base backup). 2. **ArgoCD GitOps chart bumps need hard refresh** — changing `targetRevision` in the Application manifest and pushing doesn't guarantee ArgoCD pulls the new chart. Annotate with `argocd.argoproj.io/refresh=hard` to force. 3. **Ceph CSI RBD DaemonSet upgrades disrupt API servers** — the nodeplugin DaemonSet rolls across ALL nodes including CP nodes, briefly disrupting kubelet and making API servers unreachable for 2-5 minutes. Schedule during maintenance windows. 4. **ArgoCD v2→v3 chart requires redis.networkPolicy.create** — `--reuse-values` fails with nil pointer error. Must provide explicit values file with `redis.networkPolicy.create: false`. 5. **ExternalSecret perpetual OutOfSync** — ESO adds default fields to live spec not present in Git YAML. Fix by adding those fields to the Git source YAML. 6. **`.kube/cache` gets committed by `git add -A`** — always add `.kube/` to `.gitignore` before running kubectl from a repo root. 7. **Save Helm values to IaC repo** — `helm get values` → save to `epic-2-k8s/helm//values.yaml` for reproducible installs.