The problem it exists for
How I work is about the process: specs first, every claim checked, a human signing for every line. This page is about what that process keeps producing. How it’s tested is the evidence behind both. Across around a hundred Go modules and the tools built on them, the same design rules turn up again and again, and they’re there because each one was learned the hard way at least once.
If you’re deciding whether to depend on something here, these are the promises that hold across the whole estate rather than one module at a time. They fall into six groups: how modules are grouped into families, how each one is shaped, how it reports what happened, which way its defaults lean, what a tool is handed when it runs, and how the tools treat the people running them.
Module families, not module monoliths
The most obvious pattern in the estate is the family. Wherever several vendors do the same job (chat models, git forges, config stores, message brokers, key services), the work is split into a family rather than built as one module that knows about every vendor. A small core module holds the contract: the types, the behaviour every member has to have, and the tests that prove it. Each vendor, backend or file format then gets a member module of its own, which implements that contract and is the only place that vendor’s SDK is imported.
Eight families work this way, with 46 member modules between them. The biggest is config, with 26, and chat, git and forges, messaging and signing and trust each have a showcase of their own. A ninth, go/dispatch, is being designed now, in public: a core for submitting and supervising multi-node training jobs on the scheduler a team already runs, with Slurm and Kubeflow as its first members. Its contract is already written in the same shape as the rest, with support reported three ways per cluster and a failure’s cause never collapsed into “failed” (dispatch spec 0001). Splitting things up this way is what the next three patterns grow out of.
A program carries only what it uses
Most of the toolkit started life inside go-tool-base and was lifted out into modules of its own, so a project can take the one piece it wants without the framework around it (the extraction playbook, spec 0119). Families take that further. A program links only the members it imports, so a tool that only talks to Claude never carries the OpenAI or Gemini SDKs, and a tool that reads YAML never carries a cloud SDK.
The footprint is a design decision rather than a side effect, so it’s tested. Each module carries a test that lists everything the module pulls in and fails the build if a forbidden dependency turns up: the framework itself, its config and command-line libraries, telemetry or a cloud SDK. In the words of the controls module’s own test, a regression “fails here rather than silently bloating every downstream binary”.
It also makes what a binary can do a decision made in code, not in config. Each tool lists the members it links in one file, deleting a line takes that SDK out of the build, and nothing at runtime can put it back. In some families importing the member is all it takes to switch it on (chat, forge, comms, signing and encryption), and in others you construct it yourself (config and messaging). A linked plugin is never switched on by default either, because otherwise “an import list becomes a behavioural file” (go-tool-base). Moving the adapters out of go-tool-base’s own packages took a binary that uses none of them from 81 MB to 34 MB (spec 0194).
What it costs: a member that isn’t linked fails at runtime, and a missing import is the first thing to check.
One contract, proven on every member
Because every member implements the same contract, the core can ship one shared test suite that every member runs unchanged. The suite is itself tested against deliberately broken backends, and every one of them has to fail. The forge suite’s reason: “A guard that passes everything is worse than no guard, because it certifies nothing while looking like assurance.” Messaging runs five named lies, including a backend that claims to survive a restart and doesn’t (spec 0003). Every forge and messaging member runs one, as do the chat providers (spec 0011), comms, and config’s storage backends. A new vendor proves itself by passing the same suite as the ones already there.
The fakes behind each member’s unit tests are checked against the real service too, by a suite that runs in a container when asked for. A fake “encodes those same assumptions” as the code it stands in for, and “These tests are what stop that being circular” (config-vault). How it’s tested covers both kinds of suite, and the rest of the estate’s testing. What it costs: Docker in CI for the container suites, and the container library appears in fourteen repositories’ requirement lists, though never in a consumer’s binary.
A core release reaches every member
Members aren’t pinned to each other or to a matching version. Each one releases on its own schedule, for its own fixes and features, so chat’s providers and config’s adapters all sit at different versions. What the family does guarantee is that a change to the core reaches all of them. When a core releases, the bump lands on every member as a fix, usually opened by Renovate, so colophon cuts a new release of each one, and the whole family is out on the new core together. A check on every member’s release merge request flags one that would ship against an older core (the pipeline).
What it costs: every core release becomes a release of every member, so config’s 26 adapters each release again whenever the config core does.
The shape of a module
Settings in, no configuration reading
A toolkit module takes plain settings, as a struct or as options applied over safe defaults, and the application maps its own configuration onto them. Not one module outside the config family imports config. Messaging’s limitations page gives the reason: “A toolkit module that reads configuration decides for its consumer where configuration comes from” (typed config section adapters, spec 0115). Secrets arrive the same way, as a function the module calls when it needs one, never a stored string. Secrets in use covers that half.
What it costs: every application writes the mapping.
The standard library at every boundary
Logging, HTTP and the filesystem come in as the standard library’s own types: Go’s structured logger, its HTTP client, and one common filesystem interface. A module that logs nothing when you hand it nothing never writes to the global logger. The repo module’s docs give the history: an earlier logging dependency “pulled an entire terminal-UI stack transitively” (logging through the standard library, spec 0116). No toolkit module imports a third-party logging library. The tools built on go-tool-base use a terminal-friendly backend, but behind go-tool-base’s own logger, which hands the standard library’s handler to anything else that asks.
Depend on the narrowest piece
Modules expose small interfaces for each role, and a consumer depends on the smallest one that covers what it touches. go/repo has ten, from reading a tree to tagging. Adapters declare their own one-method slice of the vendor client they wrap, so a test can drive them with a fake instead of a cloud account. What it costs: adding a method to a published interface breaks every consumer’s hand-written fake. go/repo found none of seven consumers checked their mocks against it (spec 0002).
Errors that say what to do
The estate owns its error package, go/errors, with no dependencies at all. It replaced a library that cost 46 transitive packages for the ten functions in use (spec 0001). The convention it exists to carry is that the fix goes in the error, as a hint the user sees. A forge credential that’s configured the old way says so, and points at the migration table, rather than surfacing as a 401 far from the cause. 92 Go repositories require it, and none still carries the old library.
Creating an error and reporting it are kept apart. Deep code adds the context and the hint and returns the error. It never decides whether the program should stop. One place near the top does, with the exit code travelling on the error itself, so there’s “one exit path” in the whole program (the reporting model). Errors covers creating and reporting in more depth.
It’s one of several dependencies the estate took over once their upkeep had stalled, each with the evidence written down first. colophon replaced a release tool that “has had two human commits in its last sixty” (colophon spec 0001), and yamldoc became a YAML engine of its own (spec 0008).
Honest outcomes
“Don’t know” is an answer
Capabilities are three-valued: yes, no, or unknown. A provider that can’t say whether a model supports tools reports unknown, never no, because the two-valued alternative “is a lie which suppresses working features” (chat spec 0006). Only a confident no turns a feature off, and there’s deliberately no shortcut that squashes the three into a yes or no. The same shape is used for forge sites and comms channels. What it costs: every caller handles three cases.
Bounded by default, and loss is counted
Queues, reads, caches and waits all have a limit, and where something is left unbounded on purpose, the code says why. Going over a limit is refused rather than trimmed to fit, because “a clamp changes what the job means without saying so” (afmpeg spec 0044). Whatever is dropped at a full queue is counted against the subscription it happened to. Messaging’s reason: “An unbounded queue is not a policy. It is a decision to fail once memory is gone rather than at a number somebody chose” (full queues). Limits like these are in 25 Go modules, from messaging and NATS to the HTTP client, the rate limiter and every capped read.
What it costs: swapping a messaging backend can change what happens at a full queue, because each backend can only drop messages the ways it supports. And a call that can’t be cancelled, such as a keychain write, is bounded by giving up on it, so it may still finish after you’ve moved on.
Partial success says what it dropped
When part of a request can’t be honoured, you still get a working result, along with an error listing exactly what wasn’t applied. chat’s constructor “returns the best functional client it can build, alongside an error describing everything it could not apply” (chat spec 0007). The forge adapters do the same for a label that didn’t stick, and their error “never travels alone”, so finding it is a complete test for “this partly succeeded”. Messaging falls back to a backend’s own full-queue behaviour and says so, and the schema guard lets a message it can’t check through, but counts it (failing open). Anything the module is responsible for itself still fails outright.
What it costs: Go’s habit is to ignore the result whenever there’s an error, and a caller who does that gets the old all-or-nothing behaviour. The better one is there for callers who ask for it.
Silence is a failure
A green run has to mean the work happened. A skipped test, a tag whose release never appeared, a message delivered to nobody: each is turned into something visible, because “an absence does not show up in a log” (a skip is not a pass). When one project’s skipped tests started running, its coverage went from 52% to 91%. colophon warns about a tag with no release (spec 0018), and the message bus counts a publish nobody received.
What it costs: time and noise. That same project’s tests went from 1.8 seconds to 12.3, and making image tool bumps visible meant more tags and changelog entries across ten repositories (cicd spec 0081).
Never cache a failure
Clients for AWS, Azure, Google Cloud and Vault are built once and shared, through one helper that keeps a success and forgets a failure, so the next call tries again. Caching the failure would turn a passing network blip into an outage that lasts until the program restarts, which is why Go’s own build-once helper “is explicitly not an acceptable implementation of this contract anywhere in the estate” (never cache a failure). 17 modules use it. What it costs: while a provider is down, every call pays for another attempt.
Safe by default
Opt out of safety, never in
The constructor with no options is the hardened one. The HTTP client insists on TLS 1.2 or later, caps its connection pool and refuses a redirect from https to http, and “You opt out of these … never in” (hardened defaults). transit retries only requests that are safe to repeat, because “the safe direction has to be the quiet one” (safe defaults). forge sends a credential only to the host it was configured for.
The same goes for anything that reaches a person or leaves the machine. Telemetry, posting, mentions, uploads and AI providers are all off until asked for. afmpeg caps a job’s memory and time out of the box, because “A library whose safe configuration is opt-in gets deployed unsafely” (safe defaults), and krites sends no photograph anywhere unasked, since “A default that silently uploaded photographs would be a betrayal, not a convenience” (why providers are opt-in).
What it costs: the safe default sometimes gets in the way. The HTTP client speaks HTTP/1.1 only, and because its list of key-exchange curves is pinned, it doesn’t offer the hybrid post-quantum ones Go adds to its own defaults. An afmpeg job that needs more than 512 MB fails until someone raises the limit.
Refuse, rather than ignore or guess
A setting the module can’t honour safely is refused when it’s built or loaded, with an error that names it. It’s never dropped or guessed at. The embedded NATS server won’t listen on a network until it has TLS and authentication, because “the unsafe surface does not exist rather than existing with a warning next to it”. The NATS client refuses a URL with a password in it, and refuses to send a token over an unencrypted connection. Config refuses to write a value from a secrets backend into a plain file, and refuses an HCL file that uses variables or functions. Signing refuses a key format it would otherwise have misread. Code that parses input from outside is fuzzed for the property it protects, not just for crashes: one encryption target checks that bytes appended to a published certificate can never change the key a report gets encrypted to (go/encryption).
When a decision about trust or visibility has to be made without an answer, the restrictive branch wins. A forge that can’t tell whether a repository is public reports unknown, and callers treat unknown as not public, because there a wrong answer “is a disclosure, not a bug” (forge spec 0003). There’s one deliberate exception. Self-update goes ahead on the key built into the binary when the published copy can’t be reached, so a tool behind a broken proxy can still patch itself, though two copies that disagree always stop it. krites’ model downloads refuse in the same situation, because for them an unreachable key host “means try again later”. The stakes decide.
What it costs: capability withheld until it’s safe, which is a real gap for anyone who wanted it now.
Lifecycles that can’t be revived
Two rules from controls. A registered service is required, so one that’s unready makes the whole program unready, while a worker attached to a supervisor is watched but not required (spec 0002). And anything that can only be used once, like a listener or a server, is rebuilt from scratch on every restart, never revived. A stopped copy refuses to be used again (spec 0004). That rule came from five real failures across three modules, and controls ships a generic wrapper for rebuilding on restart, plus an analyzer that names the line where a service holds on to something it didn’t build.
Props: one typed context for what a tool needs
Every command in a tool built on go-tool-base is handed the same thing: Props, a single object holding what a command might need from its surroundings. The name isn’t short for “properties”. A prop is the beam that stops a structure falling down (Props).
In pattern terms, Props is a strictly typed encapsulated context object: one object carrying the services a unit of work needs, passed down instead of a long list of parameters. It’s built by hand in main, so it isn’t an IoC container. Nothing is registered, resolved or built for you, and what goes in is exactly what main put there.
It’s deliberately not Go’s own context.Context. A context carries values as untyped keys, so a command that wants the logger has to look it up, cast it and hope, and a missing one only shows up when that line runs. Props has a typed field for each service, so asking for the logger gets a logger, checked by the compiler, and every dependency a tool has is listed in one struct you can read. In go-tool-base’s own words, “Context is for cancellation and deadlines, not dependency wiring” (architectural decisions), so the two travel side by side and each does one job.
It’s a deliberate “god object”, and a function that takes the whole of it would hide which services it really uses, the way a service locator does. The narrow interfaces are what stop that: a function that only logs asks for something that can log, not for all of Props. rust-tool-base reaches the same outcome with a generic application type and strongly typed config instead (spec 0008).
A filesystem you can swap
Commands never touch the disk directly. They go through a filesystem on Props, which is the real disk in production and an in-memory one under test, and can be layered, for instance a read-only disk with changes held in memory on top. The test helper builds a complete Props with no real filesystem, network, keychain or process exit behind it, so tests that use it are safe to run in parallel.
Assets that merge
A tool ships its defaults, templates and docs inside the binary, so it works the moment it’s installed. Each command registers its own embedded files into the shared asset store, and opening a file reads across all of them. A plain file is taken from whichever command registered last. A structured file (YAML, JSON, TOML, CSV, .env) is merged from every copy, so a feature can add its own settings to the shared defaults without editing anyone else’s (asset management).
Logging and telemetry that are always safe to call
The logger and the log level are on Props, and --debug moves the level for everything at once. The telemetry collector is never empty: when telemetry is off, it’s a stand-in that does nothing, so a command records an event without first checking whether anyone is listening. What that telemetry may collect depends on whose data it is. Usage analytics is “the vendor learning about users and runs on informed consent”, so it’s off until the user opts in. Observability, meaning traces, metrics and logs for a service, is “the operator instrumenting their own service”, so configuring it is consent enough (observability).
The invocation’s own streams
Input, output and whether a person is at the keyboard come from Props too, not straight from the process. No package in the framework names the process’s own standard input or output, and a test enforces that. So a test drives a real setup wizard just by handing it streams, and a tool can redirect its terminal without the framework noticing.
What’s deliberately left out
Props carries no database. go-tool-base’s decision record turns one down because “Different tools need different data stores: imposing a database pattern would pull GTB into application framework territory” (architectural decisions). An event bus, a task queue and a caching layer were turned down for the same reason: the framework is a foundation for tools, not an application framework.
So each tool picks its own storage, and the estate’s default is the user’s own files. keryx keeps its work in files committed to git, and krites keeps its records in files beside the photographs. Where a tool really needs a database, it has the one that suits it: phpbotscout, a bot that runs all the time, uses SQLite, and Scout.DM uses Postgres.
What it costs: a container this broad is easy to lean on for everything, which is why the narrow interfaces exist. And a tool that needs a database wires its own, with nothing from the framework to start from.
Tools people run
One engine, several faces
keryx and krites each have a command line, a local web studio and an MCP surface for agents, and all three sit on the same code and the same files. keryx’s studio “maps 1:1 to a CLI operation on the same files” and “does not introduce a second source of truth” (keryx spec 0011). krites requires that its studio “exposes nothing the CLI/engine can’t do” (krites spec 0003). go-tool-base’s MCP mode runs every tool call as a command of the same binary, so an agent and a person get the same behaviour.
Scout.DM takes it a step further. Its faces are roles, named in config: the chat platform connection, voice capture, the machine-facing operator API, the operator console and the web portal. One process runs all of them by default, which suits a single container. A deployment can instead start them separately, as a set of services that each run some of the roles and reach each other over a NATS message bus. It’s the same binary either way, because “a role boundary that exists only in production is a role boundary nobody tests”. A role name it doesn’t recognise is refused at start, since a process that starts and runs nothing looks exactly like a healthy idle one.
What it costs: the studio is built into the binary by a generate step, so a plain go install of krites ships a placeholder in its place.
Irreversible verbs stay human
Agents get the commands that read, draft and generate. Anything that publishes, signs, spends or touches credentials is kept off the MCP surface “so an assistant cannot post publicly or rotate tokens” (keryx spec 0055). sigillum keeps signing back for the same reason: “a signature, once emitted, is a claim the project cannot retract” (sigillum spec 0003).
Where something runs unattended, it acts only on an approval a person has recorded in git. keryx’s scheduled posts send only what was approved, because “Approval is the point where a person takes responsibility” (keryx concepts). phpbotscout starts on every server in shadow mode, where “Everything runs; nothing posts” (phpbotscout spec 0010).
What it costs: a person stays in the loop for every one of those steps, and phpbotscout posts only after a person has read 30 of its answers and agreed they’re worth posting. Even then the software doesn’t switch itself on, because “nothing here lifts shadow mode” (phpbotscout). A person does.
Ask the world, and a re-run finishes the job
A tool’s own files are the default store (what’s deliberately left out covers why, and the exceptions). Where a tool needs to know what happened, it asks the thing itself rather than a record of it. colophon takes versions from git, “the only source that cannot disagree with reality”, after finding 9 of 231 recent tags in the Go group weren’t on their default branch (what actually landed). skillup is built on the same idea. It takes a plugin’s previous version from the history of the plugin’s own manifest rather than from the last tag, because “A marker is a second source of truth, and a second source of truth is something that can disagree with the first” (why not tags).
That makes a re-run safe. colophon “finishes the release it started instead of adding a second one beside it, and says which of the two happened” (colophon spec 0014).
What it costs: a read before every write. And keryx keeps large media in object storage rather than git, which handles it badly.
A preview that writes nothing
Deciding, writing and pushing are separate commands, and the one that decides writes nothing and needs no credentials. colophon’s plan “Writes nothing: no branch, no merge request, no tag. Needs no credentials, so it is safe to run against any checkout” (colophon). skillup’s plan changes nothing and its apply writes one file without committing, keryx has a dry run on its writing commands, and config plans a change before applying it. What it costs: more commands to learn.
The irreversible step goes last
Every step that can fail runs before the one that can’t be undone. rust-tool-base’s updater downloads, checks the signature and checksum, unpacks and runs the new binary’s self-test before swapping it in, so “the disk holds either the old version or the fully verified new one”. colophon creates the release last, because a release with nothing on it “answers yes to the only question a consumer knows how to ask”, and two live breakages were recorded before that was fixed (complete releases).
And once released, nothing changes: “A broken release is corrected by cutting a new tag” (colophon spec 0025). What it costs: a mistake stays published, so the fix is always another release.
Decisions and what they cost
- Own the error package. Replacing a one-maintainer dependency meant migrating every module, and the saving only lands when the last one moves. The spec says so: a partly migrated estate carries both.
- Families, not a monolith. The extraction playbook’s own verdict is that moving the tests is the bulk of the work, and a family of modules means a release round across every member whenever its core changes, which a monolith never needed.
- Generality, proven first. Messaging specified three backends unlike NATS before building NATS, and they forced nine changes to the contract (messaging). That was slower, and it’s why the abstraction holds.
- The stakes set the default. Fail closed where a wrong answer leaks something or runs untrusted code, and fail open only where refusing would strand a user, as self-update does. What it cost: the estate is deliberately inconsistent here, and each exception has to be written down.
Proof in use
- 98 Go repositories, 54 of them in eight families, and every one is pre-1.0. A breaking change ships as a minor version, with the migration written into the release notes before it merges (colophon).
- go/errors is required by 92 of them, and 34 of the toolkit modules take the standard logger.
- The tool patterns hold across keryx, krites, colophon, sigillum, skillup and go-tool-base, and Rust’s framework reaches the same outcomes with its own idioms.
- The patterns are written down for agents as much as for people. The estate’s Go module skills teach them to every session that creates or uses a module (Agent skills).
Borrow it when, and when not to
Borrow these if you’re building a library other people will import, and want them to take only what they need. The dependency-budget test and the lying-backend suite carry over to any language with a dependency lister and a test runner. The tool patterns carry over to any command-line tool that acts on someone’s behalf, and most of all to one an agent will drive.
They cost more up front than a single package does: more modules, more release coordination, and a mapping layer in every application. For a single application that will never be imported, most of this is overhead.
Last reviewed .
