Showcase

How it's tested

Tests that have been watched failing, suites that prove themselves against backends that lie, and checks for the things a unit test can't see.

The problem it exists for

When agents write a lot of the code, a green test suite is easy to come by and hard to believe. A test written to pass passes, a suite that skips everything is green, and a check that has never failed may not check anything. So the estate’s testing is built around one question: has anyone seen this fail? In the words of the estate’s own test-first skill, “A test you never saw fail has not been shown to test anything.”

How I work covers the process and Estate architectural patterns the design. This page is the evidence behind both.

How it works

Red before green

A change starts with a failing test, and a bug fix starts by proving the bug against the untouched code. The estate’s skills hold every agent session to that: write the test, “watch it fail”, then write only enough code to pass it, and when something breaks, “You do not get to theorise until you have a command that goes red on this bug” (test-first-discipline). Prove the bug before you fix it is the story of why.

Every Go repository under the race detector

97 Go repositories run their tests through one cicd component, with Go’s race detector on, a coverage report and a summary of what was skipped (go-test). Its time limit is deliberately twice Go’s default, because a race-enabled suite near the limit “fails with a goroutine dump that reads exactly like a deadlock and is not one”. go-tool-base bans the package-level mocking hooks that would race under parallel tests.

Suites that prove themselves

Every family’s core ships one test suite that each member runs unchanged: 14 repositories run config’s backend suite, and the chat providers, forge adapters and messaging backends each run theirs. A suite like that is only worth something if it can fail, so messaging, forge and comms test the suite itself against deliberately broken backends. Messaging’s reason: “A capability that lies is worse than one that is absent”. Each broken backend runs in a child process, and the test “passes here only if it failed there” (messaging). chat and config check theirs against a simple in-memory provider that’s correct by construction, and none of the suites brings in an assertion library, so a provider that runs one takes on no new dependency.

A skip is not a pass

Tests that need a real service, a network or a built binary are always compiled and skipped unless an environment variable asks for them, never hidden behind a build tag, where they “rot silently” (env-gated tests). 37 repositories gate suites this way, and 14 of them start the real service in a container. A skip still has to tell the truth. When the variable is set and the dependency is missing, the test fails, because “Somebody asked for them, and they cannot run” (a skip is not a pass). That rule came from a green pipeline whose skipped tests, once running, took one package’s coverage from 52% to 91%.

Fuzzing for properties

73 fuzz targets across 12 repositories, most of them named for the property they protect rather than for “doesn’t crash”: a decryption that never emits unauthenticated text, a routing table that can never bypass discovery, a render that always keeps the body. The seeds replay in every ordinary test run. Four repositories also run a short fuzz search on every merge request, because four review rounds, “roughly 11M tokens of agent time”, never found the defect that one did: “A fuzz target found it in three seconds” (go/encryption). In yamldoc, which has 40 of the targets, “Fuzzing is a merge gate, not a nightly job”.

Properties over snapshots

A snapshot test compares output byte for byte with a saved copy, and the estate mostly avoids them. ffmpeg-wasi checks properties of its output instead, such as stream counts and durations, because a snapshot suite “goes red on every FFmpeg bump, which would make the instrument useless at the exact moment it is needed” (ffmpeg-wasi). Where snapshots are kept, they’re pinned to a reference on purpose: keryx’s card renderer is compared with images from the renderer it replaced, and regenerating them from the new one “would make the test agree with whatever the code now does, which is the same as deleting it”. sigillum turned a defect that came back six times into one structural test, because “Written as six separate findings they look like six unrelated bugs. Written as a property they are one.”

Mutation as a ratchet

Mutation testing changes the code on purpose and checks that a test notices. sigillum runs it on five packages on every merge request and fails if a change goes unnoticed. It calls it “a ratchet, not a detector”: it was added while hunting a class of defect and didn’t find it, because “The line was covered; the state was not” (sigillum). comms-discord does it by hand. A row of its contract suite only counts as covered once its check has been watched failing under the change the row names.

Scenarios for whole workflows

Behaviour that spans a whole command or a lifecycle is written as Gherkin scenarios and run against a built binary with godog. go-tool-base uses them “strategically for CLI workflows and state machine scenarios”, with table-driven unit tests as the baseline (architectural decisions), and in go-tool-base a new command needs scenarios before it merges. Ten repositories have them: go-tool-base has 290, keryx 121, colophon 97, Scout.DM 65 and krites 55. The Rust crates use the cucumber equivalent.

What a unit test can’t see

  • A dependency budget. 86 repositories carry a test that lists everything the module pulls in and fails the build when something forbidden turns up (go/redact).
  • Reproducible images. Seven build images are built twice from scratch on every change, and the job fails if they differ, “rather than asserting that it is” (go-tools).
  • Docs against the code. krites’ docs are checked on every merge request for commands that don’t exist, by a test that has itself been watched failing (krites).
  • The right architecture. go/encryption and sigillum run their suites as 32-bit builds too, because one overflow test asserts nothing on a 64-bit machine, and “A test that only means something on one architecture has to be RUN on that architecture, or it is decoration.”

Front ends and Rust

The web front ends are tested with Vitest, testing-library and svelte-check, and the studios with Playwright. Before that runs, a hard check confirms the browsers are actually installed, because the earlier job “failed in 12 seconds and allow_failure: true reported that as a green pipeline for three runs” (svelte-test). All 12 Rust repositories run cargo-nextest, because some tests “only pass with nextest’s process-per-test isolation”, alongside clippy with warnings as errors, cargo-deny and a coverage floor that fails the build (rust-test).

Decisions and what they cost

  • Runtime gates, not build tags. What it cost: container suites need a privileged runner, and a run with the variable off reports a lot of skips.
  • The race detector on every run. What it cost: a slower suite, and a real hang takes the full 20 minutes to be reported.
  • Fuzz on merge requests. What it cost: the budget is seconds per target, so a deep search still happens on a developer’s machine, and in sigillum “a quiet week fuzzes less than a busy one”.
  • Mutation where coverage lied. What it cost: twelve seconds a run in sigillum, and the reminder it carries: “Do not mistake a green run for evidence the suite is adequate.”
  • Two coverage rules. Rust’s floor fails the build. go-tool-base checks every package against 90% but only reports, so it “can never flake the pipeline red”. What it cost: a Go package can slip below the line until someone reads the report.

Proof in use

  • 97 Go repositories under the race detector, and 12 Rust repositories under nextest.
  • 73 fuzz targets in 12 repositories, four with a search on every merge request.
  • 732 Gherkin scenarios across ten repositories.
  • 86 repositories with a dependency budget, and seven images checked for reproducibility on every change.

Borrow it when, and when not to

Borrow it if agents write a lot of your code, or if your tests have ever been green while something was broken. The lying-backend suite, the skip that fails when it was asked to run, and the short fuzz search on merge requests carry over to any language with a test runner.

It costs pipeline time and a privileged runner for the container suites, and the habit of watching every new test fail first is slower than writing it to pass. For a script you’ll throw away, most of it is overhead.

Last reviewed .