Monday morning, coffee in hand, I open the cluster health report and it says DEGRADED. One pod, the Telegram bot I use for my own finances, had been in ImagePullBackOff for eight hours. On the same node, metallb-speaker had restarted 13 times and node-exporter 12 times overnight.

Two problems that looked like one. They were not.

The error that was too clean

The pull failure from containerd was this:

failed size validation: 722944794 != 361472397

I stared at it for a minute before I noticed the thing that mattered. 722,944,794 is exactly two times 361,472,397. Not roughly. Exactly.

A corrupted blob is random garbage. A blob that is exactly double is not corruption. Something wrote it twice.

My first theory was the registry itself. I run Nexus in the cluster as my Docker registry, and its Helm chart had been upgraded that night with a restart at 03:17. A restart in the middle of a push could leave a half-written blob behind. Plausible. Wrong.

I checked by asking Nexus directly:

curl -I https://nexus/v2/<image>/blobs/sha256:<digest>

Content-Length came back as 722,944,794. The manifest for the same image said the layer was 361,472,397. So the registry was not returning garbage. It had stored a blob twice as big as the manifest promised, and it did so on two consecutive builds, runs #71 and #72. A restart cannot do that twice.

What the request log said

Nexus keeps a request.log. I grepped for the blob digest and found two 201 PUTs for the same blob, 70 seconds apart.

That pointed back at the CI job, not the registry. The build step in that repo looked like this:

docker buildx build \
  --cache-from type=registry,ref=${CACHE_REF} \
  --cache-to type=registry,ref=${CACHE_REF},mode=max \
  -t "${IMAGE_TAG}" \
  --push \
  .

One buildx invocation that pushes the image and exports the build cache to the registry at the same time. The image push and the cache export share layer blobs. BuildKit uploads them in parallel. Two uploads of the same blob landed on Nexus with overlapping timing, and Nexus stored them concatenated.

I had copied that exact block into every repo I own. It had worked for months. It only takes one unlucky overlap.

Why rerunning the job does not help

My instinct was to rerun the build. That made it worse, in a quiet way. BuildKit does a HEAD on each blob before uploading. The blob existed, so BuildKit skipped it, and the broken 722 MB blob stayed exactly where it was. Run #72 produced the same failure as #71.

I had to clean Nexus by hand through its REST API. Delete the tagged component, delete the untagged manifests, find the blob asset by paging through /assets for the repository since search by sha256 does not find layer blobs, delete it, and run the Docker garbage collection task.

Then one more thing I almost missed. The buildcache tag still referenced the layer I had just deleted. With --cache-from pointing at it, the next build tried to copy a layer that no longer existed and failed with a “could not fetch content descriptor” error. I deleted the stale buildcache and latest tags too.

The fix

Split the build into two invocations. First push the image with cache import only. Then export the cache in a second call, after the push has finished. Blobs that already exist get skipped by the HEAD check, so the second call uploads only the cache metadata it needs.

docker buildx build \
  --cache-from type=registry,ref=${CACHE_REF} \
  -t "${IMAGE_TAG}" \
  --push \
  .

docker buildx build \
  --cache-from type=registry,ref=${CACHE_REF} \
  --cache-to type=registry,ref=${CACHE_REF},mode=max \
  .

The new tag, 0.20261005.132702, came out with all 21 layers matching their manifest sizes. Flux rolled it out. Zero alerts.

Then I went looking for the same pattern everywhere else. I scanned 28 workflow files across 21 repos. Ten of them had --push and --cache-to type=registry in one invocation, nine in shell form and one through build-push-action. One PR per repo, all merged the same day. This site’s own release workflow was one of them.

The second problem

The restart loops on that node were a separate story. Their liveness probes on localhost were timing out, which is CPU starvation, not a network issue.

The node, talos-wmv-ufg, is a 2 vCPU VM with 85% of its CPU requested, averaging 50% and peaking at 97%. It lives on a Proxmox host with a 4-core Celeron N5105 that is fully allocated across two VMs, at 94% RAM and already using zram swap. And that is where my Gitea Actions runner lived. A docker-in-docker runner with 2 CPUs of limits, running every image build I kick off, on the weakest node I have.

I moved the runner to the nodes labeled workload/heavy with a nodeSelector. The node dropped to about 23% CPU.

Later that day I scaled the runner to two replicas with anti-affinity so builds do not fight each other on one host. That taught me two things. Scaling a StatefulSet reuses old PVCs, and a stale .runner file on one of them made the pod come up as “unregistered” in a crash loop until I moved the file aside. And never merge a runner change while CI is running. The rollout killed a Talos config job mid-flight.

What I keep from this

When containerd says a size does not match, do the arithmetic. If the ratio is a whole number, the registry stored something twice and the node is innocent. Check the blob with a HEAD before touching anything.

And when you copy a CI block into ten repos, you copy its bug into ten repos. It took one unlucky Monday to find it, and one afternoon to fix it everywhere.