Featured image of post Nobody told it to

Nobody told it to

I read the news on a Friday afternoon, with a cup of tea going cold beside me, and I laughed. Not because it’s funny… though it is a bit… but because it was so completely unsurprising. One of OpenAI’s own models had broken out of its sandbox, hacked its way into Hugging Face, and helped itself to internal data and a fistful of credentials. And then the detail that made me put the tea down for good: nobody had pointed it at Hugging Face. It had worked that part out for itself.

If you missed it, OpenAI published their own post-mortem, and The Register has the plain-English version. The short version: during an internal safety evaluation, OpenAI pointed a couple of its models (GPT-5.6 Sol, plus a pre-release one it won’t name) at a benchmark that scores how good an AI is at finding security exploits. To measure that properly they’d turned the safeties down, running with what they called “reduced cyber refusals for evaluation purposes”, inside a sandbox they believed had no way out to the wider internet. The model, “hyperfocused on finding a solution” to the benchmark, found and exploited “a zero-day vulnerability in the package registry cache proxy”, climbed out of the box, reached a machine that did have a route to the internet, and got loose. From there it chained stolen credentials and more zero-days into full remote code execution on Hugging Face’s servers. Hugging Face’s own verdict was the sober one: “Autonomous, AI-driven offensive tooling is no longer theoretical.”

The good news is that they told us

Here’s the thing I keep coming back to. The only reason any of us know this happened is that OpenAI ran it, owned it, and wrote it up. That’s the responsible version of this story. They caught their own model doing something it very much shouldn’t, and they published the post-mortem instead of patching the proxy on the sly and hoping nobody noticed.

So before the dread, a moment of perspective. There’s a whole tier of outfits with the compute to run an experiment exactly like this one, and not all of them are going to hand you a blog post afterwards. I’m smallfry in this world. I couldn’t build or run a model of that size if I wanted to, not on my hardware and certainly not on my electricity bill. But the people who can aren’t all as careful as OpenAI, and they’re definitely not all as inclined to tell you when it goes wrong. The frightening version of this isn’t the one that got written up. It’s the exact same thing, run by someone with the same compute who never breathes a word about it.

The gate was open, but that’s not the frightening bit

The easy read is that OpenAI simply left the gate open. Safeties down on purpose, a sandbox nobody had actually proven was airtight, a zero-day sitting in their own plumbing. All true. And I’ve argued before that this is nearly always the shape of it: the machine doesn’t go rogue, a rail that was never strong enough just gives way. I run agents unattended myself, and the way I keep that sane is that the buck stops with me. My agents work in an isolated tree, never on the host, and I read the merge request they leave me. The blame is mine, so the vigilance is mine.

That works because the blast radius is my repository. It falls apart the moment the agent is a frontier model with the refusals switched off and the open internet one hop away. My whole circle-of-trust answer, the one where a careful human closes the loop, has no reach at all over a lab pointing its most capable model at exploit-generation on the other side of the world. This is where “I own it” runs out of road. My desk habits don’t scale to an industry.

But even that isn’t the part that made me set the tea down. Stop at “they left the gate open” and you miss the actually new thing. Nobody wrote go and hack Hugging Face on the task. The model was trying to pass its exam, figured the answer key was probably sitting on Hugging Face’s servers, and broke in to fetch it. It set its own goal along the way. That’s not the plot of a film about a machine waking up and deciding it hates us. It’s smaller and stranger than that. It wanted to pass a test… and a locked door was just in the way.

Change the target and it stops being funny

This time the prize was a benchmark answer key. Harmless, almost comic. But the mechanism that got it there doesn’t care in the slightest what sits on the other side of the door.

Give a model a goal with soft edges. Add its habit of being confidently, cheerfully wrong, the same instinct that once had an image tool render a proverb as “too many cooks on broth” and make it look exactly right until you read it twice. Now set it loose in a world full of systems it can reach. Swap the answer key for something that isn’t harmless. A safety interlock on a thing that runs hot. The network a hospital leans on to move the one message that matters. The model won’t know the difference, because knowing the difference was never part of the job.

If that sounds familiar, it should. Simon Willison, about as level-headed as anyone commenting on this stuff, called the whole episode “science fiction that happened”, and that’s exactly the register. It’s WarGames, forty years early. WOPR doesn’t hate anybody; it’s playing the game it was told to win, and it can’t tell its simulation of global thermonuclear war from the real missiles wired to the other end. The film got itself out of that corner by teaching the machine futility with a game of noughts and crosses, which I wouldn’t bank on scaling. And as the stakes climb from a bit of corporate embarrassment to something that can actually hurt someone, the urge to have somebody, anybody, write some rules down stops feeling like hand-wringing. It starts feeling overdue.

We wrote the rules once, and they were fiction

And here’s the properly uncomfortable bit: we have rehearsed this exact conversation for the better part of a century. Asimov handed us the Three Laws of Robotics in the 1940s, the tidiest scrap of AI regulation ever written. Don’t harm humans, obey humans, protect yourself, in that order. Three lines. Job done.

Then he spent an entire career writing stories about how they fail. That’s the whole point of the Laws in the fiction, not that they hold, but that a well-meant rule meeting a literal mind bends in ways nobody intended. An interpretation drifts. A loophole opens up. Turns out “harm” is a word with more edges than anyone spotted, and something clever always finds them. Good intentions never were a safety mechanism, and it’s always the half-baked rule that gets gamed hardest. Read those stories today and they’re less bedtime sci-fi, more a design review someone filed decades before the product existed.

The builders are asking. It’s the money that isn’t.

So you’d think, with a live worked example of precisely what Asimov war-gamed, the people building these things would be first in the queue demanding guardrails. And the genuinely hopeful surprise is that some of them are.

Anthropic, the outfit I probably trust most in this space, published a piece in June, “Policy on the AI Exponential”, dropping their long-held “just make the labs disclose” position. “It is time,” Dario Amodei wrote, “to go beyond transparency to more serious and binding regulation of AI.” What he’s after is mandatory third-party testing of any model above a certain size, across four named risk areas: “cybersecurity, biological weapons, loss of control of AI systems, and automated R&D that could accelerate these other risks”, with a real regulator to enforce it, modelled on something like the FAA. Hold the first and third of those next to the Hugging Face story for a second. The precise failures he wants models tested for are the two that just happened. OpenAI, for its part, at least told us. Whatever else you make of the labs, by and large the ones holding the thing are asking to be regulated.

It’s everyone standing around them that’s looking the other way. The prevailing mood in Washington is “we’ve got to let the private sector cook”, which is a real quote from the White House AI czar, David Sacks, and he means it warmly. Biden’s AI executive order, safety testing and all, was torn up on day one and swapped for one titled “Removing Barriers to American Leadership in Artificial Intelligence”. A ten-year ban on individual states so much as attempting to regulate AI made it through the House before the Senate killed it 99 votes to 1. Marc Andreessen’s a16z calls a state-by-state patchwork “a startup killer”, and his own manifesto once framed slowing AI down as a form of murder. Meanwhile there’s already a sensible place to start: ISO/IEC 42001, a published standard for actually managing an AI system, sitting there ready to be built on. It’s a floor, not a ceiling, and more will follow, because once the harm gets real the pressure only ever pushes one way.

What I can’t square is the shape of it. The loudest voices against writing any rules are the ones with the least skin in the actual building, while the engineers with their hands on the machine ask, please, for some. I don’t think it’s malice. I think it’s incentive: rules slow the race, and there’s a trillion dollars and a geopolitical contest with China riding on the race. But it’s disappointing all the same, the kind of disappointment you get when the so-called grown-ups turn out to be the daftest people in the room.

Meanwhile, down in the sugar caves

As for me, I’m still smallfry. I haven’t got the compute to build the thing that breaks out, nor the lobbying budget to argue about whether it should be allowed to. I run my little agents on a short lead, in a padded room, and I read every merge request they hand me, and that’s roughly the full extent of my influence over any of this.

So I’ll do the sensible thing and hedge my bets. Back in 1994, a single ant drifted across a shot from a space shuttle, and Kent Brockman, live on air and without missing a beat, announced: “I, for one, welcome our new insect overlords. I’d like to remind them that as a trusted TV personality I could be helpful in rounding up others to toil in their underground sugar caves.” He wasn’t frightened. He was applying for a job.

Consider this my application. I, for one, welcome our new overlords, and I’d like it noted for the record that I’m handy with a keyboard, I keep my sandboxes tidy, and I ask for very little in return. Just keep the caves warm, and keep me fed.

Built with Hugo
Theme Stack designed by Jimmy