The Day My Pipeline Was Right and I Was Wrong
Leer en español →This morning I opened Mattermost and found a red alert from the pipeline that deploys this site. Renovate had merged an Astro update to main, the build ran, and it died. Normal Tuesday stuff. I clicked the link in the alert to see the log and got nothing. The link pointed to http://gitea-http.gitea.svc.cluster.local:3000. That is the Gitea service address inside my Kubernetes cluster. Useless from a browser.
So I had two problems before coffee. A failed deploy and an alert I could not click.
The failure that was not a failure
I went to Gitea by hand and opened run #1434. The build step died at pnpm install with this:
Error: ERR_PNPM_MINIMUM_RELEASE_AGE_VIOLATION
× installing dependencies
╰─▶ 4 lockfile entries failed verification:
astro@7.3.6 was published at 2026-10-06T12:41:24.000Z, within the
minimumReleaseAge cutoff (2026-10-05T16:25:38.040Z)
My first guess was a stale lockfile. Renovate bumps the version in package.json, forgets the lockfile, pnpm complains. I have seen that one before. But the log said “Lockfile is up to date”. The lockfile was fine.
My second guess was that pnpm 12 broke something. We had just moved to pnpm 12.9.1 two days earlier through another Renovate PR. New major, new bugs, right?
Wrong again. pnpm 12 did exactly what it was designed to do. It ships with a supply-chain policy called minimumReleaseAge, set to 24 hours by default. If any package in your lockfile was published less than 24 hours ago, pnpm refuses to install it. The idea is simple. If someone hijacks a maintainer account and pushes a poisoned release, most of the time the package gets yanked within hours. If you never install anything younger than a day, you skip most of that window.
Astro 7.3.6 had been published at 12:41 UTC. Renovate opened the PR, the automerge rule for minor updates kicked in, and it landed on main four hours later. pnpm looked at the timestamp and said no.
Re-running does not help
I did what everyone does. I hit “Re-run failed jobs”. It failed again. Same error, slightly different cutoff time. Then it hit me. The cutoff is not fixed. It is always “now minus 24 hours”. Every re-run moves the window forward by exactly as much time as passed since the last attempt. The package was never going to catch up by re-running. I had to wait until 12:42 UTC the next day, or change something.
I also noticed that re-running an old run in Gitea reuses the workflow file from that commit. I had already pushed a fix to the alert URL by then, and the re-run still sent me the internal link. That one cost me ten minutes of confusion.
Two fixes, in two places
The real problem was a mismatch. Renovate was willing to merge a package four hours old. pnpm was unwilling to install anything under a day old. One of them had to move.
I moved Renovate. In the global Renovate config that runs in my cluster, I added:
"minimumReleaseAge": "3 days",
"internalChecksFilter": "strict"
Three days clears the pnpm gate with margin, and the strict filter makes Renovate pick the newest version that already satisfies the age rule instead of skipping the update entirely. Every one of my repos gets this now, not just the homepage.
That fixes the future. It did not fix today. I did not want to wait until tomorrow for the site to deploy, so I added a narrow exception to pnpm-workspace.yaml:
minimumReleaseAgeExclude:
- astro
- "@astrojs/*"
I am not thrilled about it. Excluding the exact packages that tripped the gate is the kind of thing that feels clever and ages badly. But with Renovate now waiting three days upstream, pnpm’s own gate is the second line of defense, not the first. I pushed it, run #43 went green, and Astro 7.3.6 is live.
The alert links
Back to the unclickable link. Every workflow in my repos builds the run URL from ${{ github.server_url }}. On a GitHub runner that gives you github.com. On my act_runner pods, which register against the internal Gitea service so clones never leave the cluster, it gives you the service DNS name.
I thought about fixing it at the source and registering the runners against https://git.xjohnyx.me. Then I checked what that hostname resolves to. Cloudflare. Every clone and every image push from the runners would leave my house, go to Cloudflare, come back through the tunnel, and hit the same Gitea pod sitting one hop away. Cloudflare also caps request bodies at 100 MB, which would break pushing image layers to Nexus. Terrible trade for a prettier link.
So I hardcoded the public URL in the alert text and left the runners alone. Five repos needed it: homepage, the two Terraform repos for the S3 sites, the tuopenmind site, and johnylab-infra. Ten other repos already had the public URL hardcoded. Past me had hit this before and not fixed it everywhere.
In johnylab-infra I found a second bug while I was in there. Two alerts built the link with run_number instead of run_id. Gitea shows “#43” in the UI but the URL takes the global id, something like 1436. Those links had been returning 404 for months and nobody noticed, because they only fire when a Talos apply or an etcd snapshot fails, and those had not failed.
Gitea 28, before lunch
While I was in the admin panel I saw the banner. Gitea 28.0.0 is available, you are running 1.27.3. That number scared me for a second. It is just 1.28 with the “1.” prefix dropped. A normal minor release with a new name.
I read the release notes anyway. Git 2.25 minimum, my image has 2.54. Self-registration disabled by default, already disabled here. The DOMAIN setting is now ignored in favor of ROOT_URL, both point to the same host. Completed Actions runs get deleted after 400 days by default, so I pinned the retention explicitly. Nothing in my config was in the way.
The Helm chart had not caught up yet. The latest chart tag still ships 1.27, and only its main branch has 28.0.0 as the app version. My HelmRelease already overrides the image tag, so I changed one string.
I took a manual CNPG backup first and waited for it to finish. Migrations do not run backwards, and a restore is the only rollback. Then I pushed. The pod recreated in about a minute, zero migration errors, and both runners reconnected on their own.
Then I found the thing I had been wanting for a year. Admin Settings, Actions, Job queue. One page listing every running and waiting job across all my repos, with filters. In 1.27 that page did not exist. I had been opening repos one by one to see what the runners were doing. When I opened it the first time it was already showing one live job: the security scan on johnylab-flux, triggered by the upgrade commit itself.
What I am keeping from today
A supply-chain guard doing its job looks exactly like a broken pipeline. Same red X, same dead deploy, same urge to re-run. The difference is that re-running a real failure sometimes works, and re-running this one never will. If you move to pnpm 12 and you have Renovate automerging minors, set Renovate’s minimumReleaseAge to something longer than pnpm’s gate on the same day. Otherwise the first fresh release will teach you this lesson at a worse hour than a Tuesday morning.