The problem it exists for
One person can’t hand-maintain a CI file in 146 repositories. Every lint rule, security scan, release step and docs deploy would be written dozens of times and drift dozens of ways. So there’s one library of CI components instead, and 133 of those projects include it: 717 include lines between them, with the release component in 117 projects and the Go test, lint and security jobs in 100 each.
And the images those jobs run on used to be heavy. The lightweight jobs were pulling an image of nearly a gigabyte, 2.72 GB unpacked, carrying 2,582 scanner findings of every severity. Today they pull ci-base, under 80 MB.
How it works
Components, not copies
cicd is one GitLab project holding 38 CI components across nine tracks (Go, Rust, OpenTofu, docs, release and the rest), all released under one version stream. A project includes the components it needs and gets the whole job, rules and caching included. Each component has a self-test pipeline of its own, so a change is proven in cicd before it reaches anyone.
Hardened images, one per job
The jobs run on a family of small images, one per track, built from a Wolfi base and pinned to an exact digest. A job pulls the image for its language and nothing else. Build images covers how they’re built, scanned and kept current.
Runners that scale to nothing
The jobs run on a runner fleet of my own: one tiny manager instance always on, and spot-priced workers started on demand, from none up to four by default. Each worker takes one job, disappears after ten idle minutes, and is replaced after a hundred jobs, so a quiet evening costs next to nothing.
Who can fix it decides the gate
A check fails the pipeline only when the people running that pipeline can fix what it finds. The image scan is split in two for that reason, “because two different people can act on them”: findings in the operating system packages fail the build, since a rebuild clears them, while findings inside upstream tools are reported for their maintainers (image-scan). The skills gate works the same way. Hidden characters in a skill fail outright, and a suspicious-looking instruction is flagged for a person to read. Where a check is advisory, the component says why in its own header, because a job allowed to fail also hides a job that’s broken.
Security runs on every change
Most jobs skip a merge request that can’t have affected them, so a docs change doesn’t run the Go tests. The security jobs never skip, because “A dependency’s vulnerability status isn’t a function of the diff”, it’s a function of time (security, always on). A docs-only change merged the day after a vulnerability is published is exactly the one the scanner should see. The one exception came from measuring it: re-scanning colophon’s release merge requests, which change only the changelog, found nothing in 790 pipelines over three weeks and cost a fifth of the estate’s merge-request minutes.
Finishing what the tools start
Several components exist to carry one of the estate’s tools into every pipeline, so the tool does the work and the component makes it happen on time, everywhere:
- colophon runs colophon’s release jobs in 118 projects, and goreleaser builds and publishes the binaries of the eight tools that ship them, go-tool-base and the tools built on it.
- go-core-currency checks, on 98 projects’ release merge requests, that each family member is built against its core’s latest release, reading the families from this site’s own projects feed. That’s what keeps families like config, chat and git and forges in step.
- skill-security is the security gate on the agent skills marketplace, and docs-verify runs a project’s own docs check, docscheck in krites’ case, on every merge request.
- go-singleuse runs the controls analyzer for restarts that would reuse something they can’t, and tofu-module-publish publishes the AWS account foundation modules to GitLab’s module registry.
Decisions and what they cost
- One library. Every component lives in one project (spec 0001). What it cost: every release reaches every consumer, which is a real load on the runners when it happens several times a week.
- Track main, on purpose. Projects include the components from
maininstead of a pinned tag (spec 0079). With 106 repositories consuming cicd, 65 releases in 90 days had each been costing about a day of runner time in churn, and GitLab’s catalog and a stable branch were both turned down as the alternative. What it cost: it goes against GitLab’s own supply-chain advice, knowingly, and it puts the weight on cicd’s self-tests.
Proof in use
- 133 of the estate’s 146 active projects include cicd, and colophon releases 117 of them through it.
- cicd is at v0.55, with 83 releases, and the images release on their own (Build images).
Use it when, and when not to
Use it as a model, or borrow components from it, if you run a lot of small GitLab projects on your own and want one place to change a pipeline.
It isn’t a general-purpose product, and the limitations page says so. It’s GitLab only, with no path to GitHub Actions, there’s no Python, PHP or Java track, the OpenTofu components sign in to AWS only, sites deploy to GitLab Pages only, and minor versions can still break things before 1.0. The runner fleet runs x86 jobs only, runs its workers privileged, and was built for one account.
Where it’s going
The open issues are mostly the fleet’s: build caches that don’t like being shared by concurrent jobs, and a GitHub mirror that keeps failing.
What it's made of
- homebrew The Homebrew tap for phpboyscout tools, used to distribute the binaries that come out of the release pipeline.
- terraform-aws-gitlab-runner-fleet A GitLab Runner fleeting fleet on AWS: one always-on manager plus scale-to-zero spot workers on the docker-autoscaler executor. Built to replace cattle-ops/gitlab-runner and expose the levers it hid, including worker disk size and a shared cache layer.
The story in posts
- The builder had been archived for a year
- Nine tags that were never on the branch
- One cache key per lockfile filled my runner's disk
- Ready for human
- The fastest job is the one you never run
- A release is just another merge request
- The component that fired on every schedule
- Stop installing the same tools on every pipeline job
- The build gate these sites never had
- The secret that wasn't on my branch
Last reviewed .