The problem it exists for
Everything else on these pages was built by one person. There are 145 active projects in the estate, and AI agents do a great deal of the typing. That combination can go two ways. Either it’s a lot of plausible-looking code nobody really decided on, or it’s a lot of carefully decided code that happened to get typed quickly. The difference isn’t the agents, it’s the way of working around them, and that’s what this page is about.
It’s also the page I’d point a CTO at, because “how does one person keep 145 projects trustworthy?” is the question underneath every other showcase here.
How it works
Decide it in writing first
Nothing non-trivial gets built without a spec it implements. A spec records what done looks like, the decisions with what was rejected and what each one costs, and the questions still open, each one marked as either mine to answer or a fact an agent can go and establish. That split is what lets an agent work unattended overnight without quietly deciding things that were mine to decide. The specs live on each project’s wiki, and they’re public: go-tool-base alone has more than 200 of them. Every “decisions and what they cost” section on these showcases is lifted from them.
The reasoning stays findable
A decision is only useful if someone can find it later, so the specs are tied to the code and the docs from both ends. Code comments point at the decision that made them by number, and in go/errors’ words, “a change contradicting a D-number moves the spec first”. Across the Go modules, more than a thousand comment lines cite a spec that way. Where a decision depends on a fact nobody has, a short timeboxed experiment settles it first, and the result is kept as a dated report the spec cites. The Go module wikis alone hold 97 of them. Rejected specs stay up too, with their reasons, so the same idea isn’t argued from scratch a year later.
The docs carry the other half. Nearly every module’s documentation has a page saying what it deliberately doesn’t do, written as a decision rather than an apology: “Each of these is a decision with a name, not a gap waiting to be filled” (messaging’s limitations). Estate architectural patterns is what all of that keeps producing.
One session per project, and the tracker between them
Each project gets its own agent session, deliberately walled off so one can’t trample another’s work. When one turns up something another project needs, it doesn’t fix it in passing or try to remember it. It writes it up as an issue on that project, detailed enough that a session arriving with none of the context can act on it. Ready for human is how that came about, after I spent a while being the message bus myself and getting it wrong.
Every review is a claim, and a claim gets checked
An agent saying “fixed”, “tested” or “this doesn’t exist” is a claim, not a result. So a bug fix starts by proving the bug against the untouched code, with a failing test handed back before any change, because from the outside a real fix, a bug that never existed and a test written to pass all look the same (Prove the bug before you fix it). Reviews get the same treatment: a design worth arguing about goes to a panel of models from different vendors, told not to agree with each other for the sake of it, and what they find is evidence to weigh rather than a verdict to accept. How it’s tested covers what that looks like across the estate’s suites.
A human signs for every line
AI usage isn’t hidden anywhere here, and it isn’t credited in the commits either, because a person is accountable for every line that gets merged and an attribution trailer blurs who that is. Agents don’t cut releases. A release is a merge request that I merge, and on the infrastructure side even an apply waits on a human merging a plan that’s already been checked (Reviewed, then applied).
The habits, packaged
All of this lives as skills that any agent session can load, in a public plugin marketplace: 66 skills across 8 plugins at the last count, from spec-driven development and proving a bug first to timeboxing and running a review panel. Because a skill is effectively source code an agent runs, the marketplace has a security gate of its own on every change (the post on why).
Decisions and what they cost
- Specs before code. Every non-trivial change has a written decision record first. What it cost: time up front on every feature, and a wiki per project to keep tidy, in exchange for agents that build what was decided instead of what sounded plausible.
- Use, contribute, then build. I’d rather use somebody else’s tool, and try to fix it upstream when it nearly fits, before writing my own (Building it yourself is the third thing I try). What it cost: a lot of the estate is still my own anyway (colophon and go/chat are both third-rung builds), and each one is something I now maintain.
- Isolation over convenience. Separate sessions per project, with the issue tracker as the only channel between them. What it cost: more issues than a team of one would normally write, and a handover has to be written as carefully as a spec.
- No AI attribution in commits. Accountability stays with the person who merges. What it cost: the agents’ part shows up in the specs and reports, where it’s credited openly, rather than in the history.
Proof in use
- 116 projects release through colophon, where merging a release is a human action by design.
- The showcases themselves: go-tool-base, chat, Signing and trust and colophon each cite the specs their decisions came from, and every one of those specs is public.
- A terrible lead to an AI junior is the counterweight: the time the way of working failed, because I treated a problem as boilerplate and briefed it badly.
Borrow it when, and when not to
Borrow it if you’re one person, or a small team, trying to get a lot done with agents without losing track of who decided what. The skills install into Claude Code, most of them carry Codex instructions too, and the general ones don’t assume my projects.
Don’t borrow all of it at once. The spec discipline pays for itself on anything you’ll still be maintaining in six months and is overkill for a weekend script, and the panel reviews cost real quota.
Where it’s going
The newest additions are a timebox that reads the real clock instead of an agent’s sense of time, and a research panel that runs one question past several models and makes them argue it out. Both shipped in October 2026, and both came out of finding the gap the hard way first.
Last reviewed .