Three Talos Hops and Four Kubernetes Minors in One Sunday
Leer en español →My homelab cluster had been sitting on Talos 1.11.0 and Kubernetes 1.34.0 for months. Seven nodes on Proxmox, three control planes, four workers, everything managed from git. I knew I was behind. What I did not know was how upgrades were actually supposed to work on my own lab, because I had never done one since the install.
So Sunday started with a stupid question: how do I upgrade this thing without breaking it? It ended with every node on Talos 1.14.2 and Kubernetes 1.37.1, zero alerts, and four script bugs I had to fix along the way. Here is what happened.
The rule I decided to follow
Talos publishes a compatibility window. Talos 1.12 supports Kubernetes up to 1.35, 1.13 up to 1.36, 1.14 up to 1.37. So the order has to be: Talos first, one minor at a time, then Kubernetes inside the window the new Talos allows. No skipping. Fresh etcd snapshot before every Talos hop.
That meant seven separate upgrades in one day:
- Talos 1.11.0 to 1.12.12, then 1.13.11, then 1.14.2
- Kubernetes 1.34.0 to 1.34.12, 1.35.9, 1.36.5, and finally 1.37.1
Each one as a PR in the infra repo, so Renovate can open the next one for me in the future instead of me noticing six months late.
Two scripts
I wrote two scripts to drive this. The first, upgrade-talos.sh, does a rolling OS upgrade to whatever installer tag is in git: apply the rendered config, drain the node, run talosctl upgrade --wait, then check the node is Ready and etcd is healthy. Control planes go first, then workers.
The second, upgrade-k8s.sh, wraps talosctl upgrade-k8s with a dry run and a step I will explain later, because it bit me.
And there was a third one I already had, drift.sh, which does a dry-run apply on every node and prints the diff between what git would render and what the node is actually running. That one turned out to be the most important tool of the day. After every hop I ran it until it said seven times “in sync”.
What broke
The generator was too new for the nodes
First mistake. I rendered configs with --talos-version 1.12 while the nodes still ran 1.11. The newer generator emitted keys like grubUseUKICmdline that old nodes reject, and it also switched to the multi-document output that does not accept my JSON patches.
The fix was to split two things I had been treating as one. The config schema version is the oldest Talos running in the cluster. The installer tag is the target. You render for what you have, you install what you want.
upgrade-k8s overwrote my CoreDNS config
After the first Kubernetes step, DNS behaved differently. Turns out talosctl upgrade-k8s rewrites the CoreDNS Corefile with its own template. Mine is managed by Flux and has a specific sequential forward policy I set after a WAN outage postmortem. The upgrade silently replaced it.
Now the script re-reconciles the Flux Kustomization that owns CoreDNS, checks that policy sequential is back in the Corefile, and restarts CoreDNS. I also turned off the --with-docs and --with-examples flags that default to true and were installing things I never asked for.
The client was older than the cluster
My laptop had talosctl 1.12.1. When I tried the Kubernetes step against 1.13 nodes it refused: “compatibility with version 1.13.11 is not supported”. Brew upgrade to 1.14.2, and I pinned that version in CI too. Simple rule I had not internalized: the client has to be at least as new as the cluster.
Fourteen PodDisruptionBudgets said no
On talos-xoq-mmc the drain hung. Fourteen PDBs with zero allowed disruptions, mostly from my CloudNativePG databases. talosctl upgrade timed out and left the node cordoned.
For a single-replica homelab this is the PDB doing its job on a situation where it cannot help. The fix was to drain workers first with kubectl drain --disable-eviction --force. Postgres still gets a graceful shutdown, and because the volumes are local-path, the pods come back to the same node anyway. Then an explicit uncordon, because the script was relying on the upgrade doing it.
I deleted the Kyverno webhook in the middle of a drain
This was the scary one. On talos-5ip-j4x the drain evicted the Kyverno admission controller. Kyverno runs with a fail-closed webhook, which means when the controller is gone the API server refuses every pod create and delete until it comes back. Including the deletes the drain was trying to do. Gitea and Nexus went down because their volumes live on that node.
The fix is ordering. Cordon, evict kyverno-admission-controller first, wait until the kyverno-svc endpoint is populated on another node, and only then drain the rest.
The afternoon problem
With everything upgraded, an alert fired: NodeSystemSaturation on talos-wmv-ufg. Load between 7 and 13 on a 2 vCPU VM, memory at 74 percent, and the Proxmox host underneath was swapping.
My first thought was a runaway pod. I looked. The biggest pod on the node used about 0.06 cores. No hog.
The real cause was the drains. During the rolling upgrade that node was repeatedly the only uncordoned worker, so every evicted pod landed on it. Thirty-nine pods, twenty-five of them in an eleven-minute window. Meanwhile talos-5ip-j4x sat at 10 percent CPU. Kubernetes schedules once and never rebalances on its own.
I fixed it by hand first. Cordon, delete 18 stateless pods that had arrived that day, uncordon. Load went from 13.5 to under 1 in five minutes.
Then the permanent fix: kube-descheduler as a CronJob every 15 minutes, using the LowNodeUtilization strategy on actual usage from metrics-server, not on requests. That distinction matters. By requests the node looked 85 percent full and no other node looked underused, so a requests-based descheduler would have done nothing. By real CPU it was obvious.
The first run taught me one more thing. It evicted a duplicate replica of this very website, and the scheduler sent it straight back to the same node, because the deployment carried a preferred node affinity for that host from a long time ago. An evict and reschedule loop. I replaced it with a pod anti-affinity on hostname, and narrowed the descheduler to only enforce required affinities, since five other apps still carry the same stale preference.
Where it ended
Seven of seven nodes on Talos 1.14.2 and Kubernetes 1.37.1. Drift script green seven times. etcd defragmented from 262 MB to 128 MB. CNPG eleven of eleven healthy. Zero alerts, zero open PRs.
What is still open: Talos 1.14 generates multi-document machine configs, and every document conflicts with my old v1alpha1 patches. The nodes run fine on the legacy layout, so I pinned the schema at 1.13.11 and wrote the migration down as debt. That is the next Sunday.