230 lines
8.4 KiB
Markdown
230 lines
8.4 KiB
Markdown
# Helm Chart Upgrades — 2026-07-12
|
||
|
||
## Goal
|
||
|
||
Upgrade all Helm-managed charts on the RKE2 cluster after the v1.35.6
|
||
upgrade. 5 components needed updates; RKE2-bundled charts were already
|
||
current.
|
||
|
||
## Version Matrix (Final)
|
||
|
||
| Component | Old Version | New Version | Method | Status |
|
||
|-----------|------------|-------------|--------|--------|
|
||
| External Secrets | 2.2.0 | 2.7.0 | Helm direct | ✅ |
|
||
| CNPG | 0.22.1 (app 1.24.1) | 0.29.0 (app 1.30.0) | ArgoCD GitOps | ✅ |
|
||
| Velero | 8.1.0 (app 1.15.0) | 12.1.0 (app 1.18.1) | ArgoCD GitOps | ✅ |
|
||
| Ceph CSI RBD | 3.10.1 | 3.17.0 | Helm direct | ✅ |
|
||
| ArgoCD | 7.8.13 (app v2.14.7) | 10.1.3 (app v3.4.5) | Helm direct | ✅ |
|
||
| Cilium | rke2-cilium-1.19.402 | — | RKE2-bundled | Already current |
|
||
| Traefik | rke2-traefik-40.1.003 | — | RKE2-bundled | Already current |
|
||
| CoreDNS | rke2-coredns-1.46.002 | — | RKE2-bundled | Already current |
|
||
| Metrics Server | rke2-metrics-server-3.13.100 | — | RKE2-bundled | Already current |
|
||
| Snapshot Ctrl | rke2-snapshot-controller-4.2.006 | — | RKE2-bundled | Already current |
|
||
|
||
## Upgrade Order (Safest → Riskiest)
|
||
|
||
1. External Secrets 2.2.0 → 2.7.0 (Minor, Helm direct, minimal values)
|
||
2. CNPG 0.22.1 → 0.29.0 (Major, ArgoCD GitOps, **required replica fix first**)
|
||
3. Velero 8.1.0 → 12.1.0 (Major, ArgoCD GitOps)
|
||
4. Ceph CSI RBD 3.10.1 → 3.17.0 (Major, Helm direct, **causes transient API outage**)
|
||
5. ArgoCD v2.14.7 → v3.4.5 (MAJOR v2→v3, Helm direct, **breaking change**)
|
||
|
||
## Detail: External Secrets (2.2.0 → 2.7.0)
|
||
|
||
Straightforward `helm upgrade --reuse-values`. Values were minimal
|
||
(`installCRDs: true`). No breaking changes. All ExternalSecrets
|
||
continued syncing. Two pre-existing broken ExternalSecrets
|
||
(argocd-repo-credentials, postgres-main-db) were already broken for
|
||
110+ days — unrelated to the upgrade.
|
||
|
||
## Detail: CNPG (0.22.1 → 0.29.0) — Replica Rebuild Required
|
||
|
||
### Pre-Existing Issue Found
|
||
|
||
`postgres-main-2` replica was broken for 109 days:
|
||
```
|
||
requested WAL segment 000000020000000000000008 has already been removed
|
||
WAL file not found in the recovery object store
|
||
```
|
||
Primary had recycled WAL segments the replica still needed. Cluster
|
||
was running with 2/3 Ready (primary + 1 replica), backups operational.
|
||
|
||
### Replica Rebuild
|
||
|
||
```bash
|
||
# Delete broken replica PVCs + pod
|
||
kubectl delete pvc postgres-main-2 postgres-main-2-wal -n postgres
|
||
kubectl delete pod postgres-main-2 -n postgres
|
||
|
||
# CNPG automatically created postgres-main-4 as replacement
|
||
# Wait for 3/3 Ready:
|
||
kubectl get pods -n postgres -w
|
||
```
|
||
|
||
### Operator Upgrade via GitOps
|
||
|
||
```bash
|
||
# Changed targetRevision in ArgoCD Application manifest
|
||
cd /root/iac-homelab
|
||
sed -i 's/targetRevision: 0.22.1/targetRevision: 0.29.0/' \
|
||
clusters/main/operators/cloudnativepg.yaml
|
||
git add -A && git commit -m "feat: upgrade CNPG 0.22.1 → 0.29.0" && git push
|
||
```
|
||
|
||
**Pitfall**: ArgoCD didn't auto-detect the chart version change despite
|
||
`selfHeal: true`. Required hard refresh:
|
||
```bash
|
||
kubectl annotate application cloudnativepg-operator -n argocd \
|
||
argocd.argoproj.io/refresh=hard --overwrite
|
||
```
|
||
|
||
After refresh, ArgoCD pulled new chart, created new operator pod with
|
||
image v1.30.0. Postgres cluster briefly went 2/3 during reconciliation
|
||
(operator triggered primary restart), then recovered to 3/3 Healthy.
|
||
|
||
## Detail: Velero (8.1.0 → 12.1.0)
|
||
|
||
Same GitOps pattern as CNPG — changed `targetRevision` in ArgoCD
|
||
Application manifest, pushed, hard-refreshed ArgoCD. All Velero
|
||
components rolled out: server + 6 node-agents (one per K8s node).
|
||
BSL (Backup Storage Location) remained Available throughout. Backup
|
||
schedules (daily-full-cluster, weekly-full-cluster) continued working.
|
||
|
||
## Detail: Ceph CSI RBD (3.10.1 → 3.17.0) — API Server Disruption
|
||
|
||
### Upgrade
|
||
|
||
```bash
|
||
helm upgrade ceph-csi-rbd cephcsi/ceph-csi-rbd \
|
||
--version 3.17.0 -n kube-system --reuse-values --timeout 120s
|
||
```
|
||
|
||
### Transient API Server Outage
|
||
|
||
The nodeplugin DaemonSet rolling update disrupted kubelet on CP nodes.
|
||
All 3 API servers became unreachable:
|
||
```
|
||
The connection to the server 10.0.30.51:6443 was refused
|
||
```
|
||
|
||
Recovery took ~5 minutes. Monitored with:
|
||
```bash
|
||
curl -sk --connect-timeout 5 https://10.0.30.51:6443/healthz
|
||
# 401 Unauthorized = server is UP (expected without credentials)
|
||
# Connection refused = server is DOWN
|
||
```
|
||
|
||
After recovery, provisioner pods (2 replicas) were 7/7 Running with
|
||
image `quay.io/cephcsi/cephcsi:v3.17.0`. One nodeplugin pod stayed in
|
||
`ContainerCreating` for several minutes (image pull delay on CP-01).
|
||
All 7 PVCs remained bound throughout — no data impact.
|
||
|
||
### Root Cause
|
||
|
||
The Ceph CSI nodeplugin DaemonSet runs on EVERY node including CP
|
||
nodes. Rolling the DaemonSet causes the CSI driver socket to be
|
||
briefly unavailable, which can block kubelet operations on CP nodes,
|
||
temporarily preventing the API server from serving requests.
|
||
|
||
## Detail: ArgoCD v2.14.7 → v3.4.5 (MAJOR Upgrade)
|
||
|
||
### Safe Path: Two-Step Upgrade
|
||
|
||
Step 1: Patch to latest v2.x (low risk):
|
||
```bash
|
||
helm upgrade argocd argo/argo-cd --version 7.9.1 -n argocd --reuse-values
|
||
# v2.14.7 → v2.14.11
|
||
```
|
||
|
||
Step 2: Major v2→v3 upgrade with explicit values:
|
||
```bash
|
||
# v3 chart requires redis.networkPolicy.create field (breaking change)
|
||
# --reuse-values FAILS with:
|
||
# nil pointer evaluating interface {}.create
|
||
# Must provide explicit values file with redis.networkPolicy.create: false
|
||
|
||
helm upgrade argocd argo/argo-cd --version 10.1.3 -n argocd \
|
||
-f /tmp/argocd-v3-values.yaml --timeout 300s
|
||
```
|
||
|
||
### v3 Values File
|
||
|
||
Key addition vs v2 values:
|
||
```yaml
|
||
redis:
|
||
networkPolicy:
|
||
create: false
|
||
```
|
||
|
||
All other values carried over unchanged (server.insecure, ingress,
|
||
replicas, affinity).
|
||
|
||
### Post-Upgrade Verification
|
||
|
||
- All ArgoCD pods Running (server ×2, repo-server ×2,
|
||
application-controller, applicationset-controller, dex-server,
|
||
notifications-controller, redis)
|
||
- Image: `quay.io/argoproj/argocd:v3.4.5`
|
||
- All applications Synced (except `backups` — pre-existing
|
||
ExternalSecret OutOfSync, see §13.7)
|
||
|
||
## Detail: ArgoCD `backups` App OutOfSync (Task 4)
|
||
|
||
### Root Cause
|
||
|
||
Single resource OutOfSync: `ExternalSecret/velero-s3-credentials` in
|
||
`velero` namespace. ESO adds default fields to the live spec that
|
||
aren't in the Git YAML:
|
||
- `conversionStrategy: "Default"`
|
||
- `decodingStrategy: "None"`
|
||
- `metadataPolicy: "None"`
|
||
- `deletionPolicy: "Retain"`
|
||
- `engineVersion: "v2"`
|
||
- `mergePolicy: "Replace"`
|
||
|
||
ArgoCD's `ignoreDifferences` only covers `/status`, not `/spec`.
|
||
|
||
### Impact
|
||
|
||
Cosmetic — ExternalSecret functions correctly (`secret synced`,
|
||
Velero backups operational). No data or functionality impact.
|
||
|
||
### Fix (Not Yet Applied — Read-Only Investigation)
|
||
|
||
Add ESO default fields to Git YAML at
|
||
`clusters/main/backups/velero/external-secret.yaml` so spec matches
|
||
exactly. Alternative: expand `ignoreDifferences` to cover specific
|
||
JSON pointers (less clean).
|
||
|
||
## GitOps Trail
|
||
|
||
Commits pushed to `dominik/iac-homelab` on Gitea:
|
||
- `36eb5c4` — `feat(helm): upgrade all managed charts` (values files +
|
||
ArgoCD Application targetRevision changes)
|
||
- `acd3231` — `chore: remove .kube cache from repo, add to .gitignore`
|
||
(cleanup of accidentally committed `.kube/cache/` directory)
|
||
|
||
## Lessons Learned
|
||
|
||
1. **Always check CNPG cluster health before upgrading the operator** —
|
||
a broken replica (WAL gap) should be rebuilt first (delete PVC + pod,
|
||
CNPG auto-creates fresh replica from base backup).
|
||
2. **ArgoCD GitOps chart bumps need hard refresh** — changing
|
||
`targetRevision` in the Application manifest and pushing doesn't
|
||
guarantee ArgoCD pulls the new chart. Annotate with
|
||
`argocd.argoproj.io/refresh=hard` to force.
|
||
3. **Ceph CSI RBD DaemonSet upgrades disrupt API servers** — the
|
||
nodeplugin DaemonSet rolls across ALL nodes including CP nodes,
|
||
briefly disrupting kubelet and making API servers unreachable for
|
||
2-5 minutes. Schedule during maintenance windows.
|
||
4. **ArgoCD v2→v3 chart requires redis.networkPolicy.create** —
|
||
`--reuse-values` fails with nil pointer error. Must provide
|
||
explicit values file with `redis.networkPolicy.create: false`.
|
||
5. **ExternalSecret perpetual OutOfSync** — ESO adds default fields to
|
||
live spec not present in Git YAML. Fix by adding those fields to the
|
||
Git source YAML.
|
||
6. **`.kube/cache` gets committed by `git add -A`** — always add
|
||
`.kube/` to `.gitignore` before running kubectl from a repo root.
|
||
7. **Save Helm values to IaC repo** — `helm get values` → save to
|
||
`epic-2-k8s/helm/<chart>/values.yaml` for reproducible installs.
|