The problem it exists for
A long-running Go program usually runs several things at once: web servers, gRPC servers, background workers. They all have to start, report whether they’re healthy, restart when they fall over, and shut down in a sensible order without hanging. Some of them are there for the life of the program, and some come and go while it runs: a worker per tenant, a subscription per topic. Most teams write that supervisor again for every project, and then write the glue that plugs each kind of server into it again too.
How it works
The controller: the services a program needs
controls has two layers. The controller looks after a fixed set of services, registered before it starts. It starts them all at once, gathers their health checks into status, liveness and readiness reports, restarts a failed service with growing waits up to a limit, and stops them in reverse order of registration within a time limit. A stop that hangs is abandoned at the deadline and named in a warning, so a shutdown always finishes. It depends on nothing but go/errors.
The supervisor: workers that come and go
The supervisor looks after children that attach and detach while the program runs, under the same restart rules, and stops them all at once. A supervisor is itself a service, so it registers with the controller like anything else: the controller supervises the supervisor, and the supervisor supervises the children. A program has one controller and as many supervisors as it has changing sets.
What separates the two layers is what a failure means (spec 0002). A registered service is required, so one that’s unready takes the whole program out of rotation. An attached child isn’t. When it fails for good, the supervisor stays ready, reports degraded, and hands the failure to the code that attached the child, which knows whether it mattered.
Common services, registered in one call
The services most programs run already know how to plug in, so wiring them up is a call rather than a page of start and stop code:
- transport registers an HTTP server, a gRPC server or a REST gateway with a controller in one call, which brings its start, stop and status with it. It also turns the controller’s health reports into ready-made health, liveness and readiness endpoints, and into the standard gRPC health service.
- messaging’s bus has the start, stop and readiness methods a controller expects, and it uses both layers: the bus is a service of the controller, and each subscription is a child of a supervisor the bus owns. A restart replaces that whole set of subscriptions at once. The bus also offers a readiness check that reports degraded, by name, when a subscription has failed for good.
- go/nats registers an embedded NATS server, and a client connected to it, with their start, stop, readiness and liveness wired up.
Decisions and what they cost
- One owner for signals. The supervisor stopped catching Ctrl-C by default: a standalone program asks for it, and the old opt-out was deleted outright (spec 0001), instead of every framework tool having to remember to switch it off. What it cost: a break for anyone relying on the old default, and a survey found nobody was. It fixed a live bug on the way.
- Registered means required. A failed child never makes its supervisor unready, at any proportion, including all of them (spec 0002). Otherwise one dead worker would reach the whole program’s health through the supervisor’s own registration. What it cost: a child that really was needed has to be noticed by the code that attached it, which then makes itself unready.
- Generations. Five recorded failures across the estate came from a restart reusing something that could only be used once. Now stopping always finishes within its limit, and an old and a new copy never run at the same time (spec 0004). What it cost: leftover work from a stopped copy may still be running, so it’s cut off and refused any further access. controls ships an analyzer that names the line where a service holds on to something a restart can’t reuse, and cicd’s go-singleuse component runs it on merge requests as an advisory check. No project includes it yet.
- Say why it stopped. A controller now records why each shutdown happened and who failed on the way, and its failure events never hold a service up (spec 0008). What it cost: a reader that stops taking events can delay shutdown by the rest of the time limit. It’s on the main branch, not yet in a release.
Proof in use
- Ten public projects run under controls, including go-tool-base, messaging, keryx, krites, phpbotscout and the transport stack.
- 17 releases of controls since July 2026, and over 200 of its own test functions.
Use it when, and when not to
Use it if a Go program runs several long-lived services, or workers that come and go, and you want them started, watched, restarted and stopped properly, without a framework around them.
It serves no health endpoints itself (transport turns its reports into them), doesn’t order startup or make one service wait for another, and a crash in a service’s start code takes the program down. A supervisor can’t be reconfigured while it runs, restart waits double with no randomness, and there are no metrics or tracing. An HTTP or gRPC server registered through transport can’t be given a restart policy yet, because a Go server can’t serve again once it has been shut down.
Where it’s going
Letting a controller stop when a service gives up, surviving a crash in a service’s start, finishing cleanly when the work is done, and one event queue for the controller and its supervisor.
Ready-made integrations
- messaging One way to send a message between services: a controls-managed bus carrying CloudEvents, with the thing that carries them behind a swappable backend.
- nats NATS as an estate convention: a controls-managed embedded server and a client that is the same code whether the broker is in this process or a cluster somebody else runs.
- transport A framework-free HTTP + gRPC + gateway server stack: hardened server constructors, health endpoints, authentication and security headers.
Last reviewed .