Featured image of post The bill for fifty-one modules

The bill for fifty-one modules

Change one line in the tls module and you've signed up for twelve releases. That's not a bug in the architecture, that's the architecture.

Change one line in the tls module and you have signed yourself up for twelve releases (releases, not commits), in four sequential waves, each one waiting on the one before it, and every single one gated on a human pressing a button. I know, because I was the human.

What the review actually called it

I had an architectural review run across the whole toolkit last July, and it came back with sixty-odd findings, the usual mix of things I already knew and things I’d rather not have. The one that mattered wasn’t a bug at all. It was theme five, and it got a title I’ve not been able to un-see:

Propagation debt as the ecosystem’s principal systemic risk. The many-small-modules architecture is clean but expensive: a tls change fans out into 12 human-gated releases across 4 sequential waves; a config change into ~20–23.

Principal systemic risk. Not the credential-timeout bug or the redirect-token hardening… the way I’d built it! At that point there were fifty-one modules.

Why it costs what it costs

The maths isn’t complicated, it’s just relentless. tls sits near the bottom, above it are the transports, above those the framework, and above the framework the tools. Change tls and nothing above it moves until tls cuts a release, then its direct consumers bump and they cut releases, then their consumers bump. Four layers deep, twelve artefacts, and each hop is a merge request I have to look at.

I’m not sure any of it counts as waste, exactly. Each of those releases is a real version of a real module that did change, and it’s the correct behaviour of a correctly built dependency graph.

It’s also most of a night, for one line.

The finding underneath the finding

What I keep coming back to isn’t the twelve so much as this:

The evidence shows propagation stalls where fan-out is largest.

It isn’t the obvious thing. You’d assume the big fan-outs get done first, on the grounds that they’re the important ones, and the opposite happens: the change with twelve downstream releases is the one you look at late on a Friday and decide is a problem for next week, and then the week after, while the modules furthest down the chain sit three minors stale and the fix waits, merged and unreleased, at the bottom.

So the debt accrues fastest where it hurts most, simply because that’s where the activation energy is highest.

What was meant to happen

The review proposed two moves worth more than the rest put together. Re-pin the whole config family in one sweep, which clears two of the ecosystem findings at a stroke. And fleet-manage the CI component pin rather than nudging repos one at a time.

It also floated something I’ve been chewing on since: whether the config adapters ought to live in one multi-module repo with synchronised tagging. That would be a partial retreat from the thing I split up to begin with, and it might still be the right call.

Where it stands now

I checked, while writing this. There are sixty-seven modules, not fifty-one… sixteen more mouths to feed than the review was even looking at. And the config family, which the review singled out as the worst fan-out in the estate, is clean: twenty-six adapters, all pinned to config v0.15.0, which is the current core, with nothing stale and nothing sat three minors back hoping I don’t look. That’s the number I expected to be bad, and it’s the one that came back spotless (it didn’t get that way on its own).

Getting the invoice down

The cost was never the extraction, because you budget for the extraction. The cost is a recurring tax on every change afterwards, and the awkward part is that it scales with how well you did the splitting. Smaller units of work, cleaner layering, a graph that actually reflects how the thing is built. And the price of all that is twelve releases for one line. Do it badly and you get a mess; do it properly and you get a bill.

So the question was never whether the architecture was worth it, it was how cheap I could make the tax. Three things, none of them finished. Renovate does the bumping, fleet-wide, from a single group-level bot rather than a config per repo. That turns “I remember to update twenty-six adapters” into something that happens whether I remember or not. It’s the reason all twenty-six are sat on the same core without me having chased a single one of them.

The pipelines do the gating. A release is a merge request like any other, so the same guards that stop a bad commit stop a bad release, and nothing needs a human except the button.

And a release-train skill, currently sat in review, which is the piece I’ve been missing. Releasing one repo is solved and has been for ages. Releasing twenty-six that depend on each other is a different problem, and it has a right answer that isn’t the obvious one. Cutting the leaves first looks like it dodges a Renovate cycle:

OrderReleases
Leaves first: 26 adapters, then core, then Renovate bumps 26, then release again53
Upstream first: core, then Renovate propagates, then adapters absorb it and release once27

Cutting the leaves first doesn’t avoid the cycle. It defers it until after the releases, so every leaf gets cut a second time the moment the core moves. Upstream-first costs you a wait while Renovate propagates, and nothing else. I had that backwards while driving the config family’s release train (for about an hour, which was long enough to feel it) and corrected myself before I doubled my own workload. That’s why it’s going into a skill rather than into my head: the mechanics are easy and the reasoning is what fails.

What I’m not paying for

There is, of course, a shop-bought answer to a chunk of this. GitLab has merge trains and they’d help, but they also live behind a subscription tier I’ve opened, looked at, and closed the tab on more than once. So I’m building my own out of Renovate, CI components and a skill that knows which way round to cut a release… on the grounds that I’d rather spend the evenings than the money. Ask me again when I’ve made my millions.

Sixty-seven, and counting

Fifty-one when the review was written, sixty-seven as I type this. That number going up is the estate working: each of those is a thing you can use without taking the rest of it, which was the point of the exercise. There’ll be another one along next week, there always is, and the trick was never to stop adding them, it’s to get to the point where the sixty-eighth doesn’t cost me a Saturday.

Built with Hugo · Theme Stack designed by Jimmy