Featured image of post The runner fleet that scales to zero

The runner fleet that scales to zero

The homelab finally ran out. Not disk this time, capacity. So the overflow went to AWS, on the condition that it costs nothing when nobody's pushing.

My homelab ran out, and for once it wasn’t disk. It was capacity: there was simply no more memory in the Proxmox box to hand to a runner, and I’d spent two weeks being clever about that instead of admitting it.

The mitigations had all landed. Capacity hadn’t moved

Most of the effort went here, and none of it made a dent. I’d serialised the greedy Rust jobs so they couldn’t take every slot at once, shared the vulnerability-database cache so twenty repos weren’t each downloading the same thing, and moved the nightly sweep to midnight where no one was waiting on it. Each of those was the right fix and each bought time, and not one of them added a single unit of capacity. That’s the thing about tuning. You can shave the nought-to-sixty all day and it does nothing whatsoever to the top speed, because the top speed is the engine. Eventually a fleet of agents raises twenty merge requests in an afternoon, twenty pipelines want to run, and three slots is three slots no matter how politely they’re scheduled.

The one condition

So: cloud. For me that comes with a condition attached, because I dislike a standing bill rather more than I dislike waiting. The requirement was scale to zero: something I could switch on when it mattered, that cost nothing at all when nobody was pushing.

And I think that fits this workload rather than being a frugal compromise. My load isn’t steady, it’s violently spiky. A morning where four agent sessions each raise five MRs, then eight hours of absolutely nothing while I sleep. A permanently running box is the wrong tool for that even when someone else is paying, because you’re renting the peak and using the mean.

GitLab has a shape for this, and it’s worth a paragraph if you’ve only ever installed a runner the ordinary way. Normally a runner is a machine. You install it somewhere, it registers itself, and it sits there waiting for work. Whether it’s busy or idle it’s switched on, and its capacity is whatever that box can do.

Fleeting splits that in two. A small always-on manager takes the jobs but never runs them. Instead it provisions a fresh throwaway worker for each one, hands the job over, and destroys the worker afterwards. The pool of those workers is the fleet, and because it’s provisioned on demand it can legitimately be empty. Less “hire a permanent employee and hope there’s enough work”, more “ring the agency when something comes in”.

terraform-aws-gitlab-runner-fleet, the OpenTofu module I built round that, makes the point in a variable description:

idle_count is fixed at 0. Pure scale-to-zero.

It isn’t configurable. There’s a knob for how long an idle worker lingers before it’s scaled in, currently ten minutes, which gives a bursty backlog a short reuse window. But the floor is zero and you can’t set it otherwise, because the moment that number can be anything else, one day it will be.

Sized by the worst job, not the average one

This is where the homelab paid for itself. The first draft of the sizing had workers on 50 GB and I rejected it on sight, because a single Rust job will eat that and ask for more, and I know that because it once pulled down 170 GB of build artefacts on hardware I owned. Workers are on 200 GB. The variable says why, in the file, so I don’t trim it later to save four quid:

Docker images + a single Rust build (cargo + target/) need generous headroom.

The manager gets 8 GB, because it stores no build artefacts at all and there’s no reason to pay for a disk that will never be written to. Disk gets sized by your worst job, not your typical one, and averages are how you end up with a fleet that works beautifully right up until it doesn’t.

Port the caching, don’t re-learn it

The other thing that came over wholesale was the caching. A fresh fleet with no shared vulnerability-database cache would have recreated, on someone else’s hardware and my credit card, the exact bottleneck the homelab had spent months teaching me about: every worker downloading the same database, every pipeline paying for it, and me wondering why the expensive thing wasn’t faster.

Those decisions had to be translated rather than rediscovered, and translating them was most of the actual work. The interesting infrastructure was about four hundred lines of OpenTofu; the rest was carrying forward everything the old runner had already taught me.

What it looks like

One always-on manager, cheap and small, and underneath it a spot worker pool with mixed instance types spread across availability zones so a spot reclaim in one doesn’t take the fleet with it. Four workers maximum, one job each, which keeps the instance ceiling and the concurrent-job ceiling the same number so there’s only one thing to reason about when it isn’t keeping up.

That ceiling is a cost guardrail rather than a performance setting, and the module says so out loud. It’ll refuse anything above 50, with an error message that reads “keep it modest and raise deliberately”, which is a note to myself at three in the morning as much as anything. The homelab runner didn’t go anywhere. It’s still there, still doing the steady work. The fleet is overflow.

Two small humiliations

The AMI shipped with a stale runner version and needed upgrading on day one. A managed image is a snapshot of someone else’s to-do list, and you inherit whatever they hadn’t got round to. And the module wouldn’t resolve from the Terraform registry. I went round the houses on URLs, module paths, naming conventions, the lot.

The project was private!

A private module in a public registry looks exactly like a broken URL, and there’s nothing in the error to tell you otherwise.

About the man with the homelab

I write about self-hosting a fair amount, and I’m the man with the Proxmox box in the corner, so this wasn’t a change of heart about cloud. It was arithmetic: the agents outgrew the hardware, and I couldn’t buy more memory for a machine that’s already full (I did look). What made it palatable was the zero more than the elasticity. I don’t mind paying for the burst. I mind paying for the silence… and there’s rather a lot of silence in a one-man estate.

The module is public and the source is here, if you fancy it. The 200 GB is not a typo.

Built with Hugo · Theme Stack designed by Jimmy