Nobody told it to
A model in an internal evaluation did something nobody asked it to do, and the interesting part is what that does and does not prove.

A model in an internal evaluation did something nobody asked it to do, and the interesting part is what that does and does not prove.

Running three different coding agents against one repository, and the house rules that stopped them undoing each other's work.

Adding stream-copy remux to a WebAssembly FFmpeg build, where writing Matroska hung forever and the identical MP4 path did not.

The kill-switch answer to AI autonomy does not survive contact with how these systems actually run. Governance is not a button.

An AI repair agent declared success the moment the code compiled. The linter disagreed, and the gap between those two is the whole story.

Hundreds of malicious package versions across three ecosystems, and the uncomfortable fact that installing one runs arbitrary code.

Building an AI dungeon master to test a Go chat client, and the bug that kept resurrecting a goblin the player had definitively killed.

Generate-and-hope does not work for codegen. Handing the draft to a repair agent with a fixed toolset that has to make build, test and lint pass.

Tool calling and the reason-act-observe loop explained, and how to define a tool in Go so the schema comes from the type rather than by hand.
