<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>PHP Boy Scout — What I think about AI</title><link>https://phpboyscout.uk/topics/thinking-about-ai/</link><description>Arguments rather than engineering: what agentic AI does to the junior pipeline, why kill switches are not governance, and what a vendor outage exposes.</description><generator>Hugo</generator><language>en-GB</language><copyright>Matt Cockayne</copyright><lastBuildDate>Sat, 01 Aug 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://phpboyscout.uk/topics/thinking-about-ai/index.xml" rel="self" type="application/rss+xml"/><item><title>The rung we sawed off</title><link>https://phpboyscout.uk/the-rung-we-sawed-off/</link><pubDate>Wed, 17 Jun 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/the-rung-we-sawed-off/</guid><category>ai</category><category>Soapbox</category><description>The junior developer pipeline is collapsing, and the cause is not the thing everyone is blaming. Who is left when the greybeards retire.</description><content:encoded>&lt;p&gt;I was in a job interview yesterday, on the wrong side of the desk for once. After
years of being the one asking the questions I&amp;rsquo;m having a look at what&amp;rsquo;s next, and
somewhere in a long, wandering technical conversation the inevitable arrived: where
do I think AI is going, and what does it mean for how we build software?&lt;/p&gt;
&lt;p&gt;I gave my answer. You can probably guess most of it. The more interesting thing was
the question I&amp;rsquo;ve started asking &lt;em&gt;them&lt;/em&gt; back. Not the salary, not the stack. What is
your actual position on AI, and how are you building a team out of both its human and
its non-human parts? I ask the company and I ask the interviewer personally, because
the two answers are rarely the same, and because I&amp;rsquo;ve decided I can&amp;rsquo;t work somewhere
that hasn&amp;rsquo;t sat with the question properly.&lt;/p&gt;
&lt;p&gt;Here is why it has become my litmus test.&lt;/p&gt;
&lt;h2 id="the-rung-and-whos-standing-on-it"&gt;The rung, and who&amp;rsquo;s standing on it
&lt;/h2&gt;&lt;p&gt;I wrote recently that
&lt;a class="link" href="https://phpboyscout.uk/the-greybeards-edge-was-never-typing/" &gt;the greybeards&amp;rsquo; edge was never typing&lt;/a&gt;:
agentic tools give a senior a boost because they have the judgement to steer and
verify, and give a junior a drag because they don&amp;rsquo;t have it yet and the machine hands
them more rope than they can hold. The cold incentive that falls out is to hire
seniors and automate the juniors.&lt;/p&gt;
&lt;p&gt;The data has since caught up with the worry. Entry-level software postings have fallen
by something like 40% from their 2022 peak. The share of juniors and graduates in IT
employment has dropped from roughly 15% to 7% in three years, and Stanford researchers
tracking early-career workers in AI-exposed jobs found the youngest cohort down sharply
from its peak.
&lt;a class="link" href="https://www.softwareseni.com/what-the-data-actually-shows-about-ai-and-junior-developer-employment-decline/" target="_blank" rel="noopener"
 &gt;The numbers are genuinely grim&lt;/a&gt;,
and plenty of people are putting it bluntly: the industry killed the junior on purpose.&lt;/p&gt;
&lt;p&gt;That framing is half right, and I think it&amp;rsquo;s worth getting the other half right too.&lt;/p&gt;
&lt;h2 id="it-was-never-about-efficiency-it-was-about-cost"&gt;It was never about efficiency. It was about cost.
&lt;/h2&gt;&lt;p&gt;We didn&amp;rsquo;t automate the junior because the work needed doing better. We did it because
people are expensive. We need sleep, we draw a salary, and our thinking takes time and
effort that a quarterly target can&amp;rsquo;t see the point of. AI got sold as round-the-clock
labour with none of that overhead, and to a business that is an almost irresistible
line on a spreadsheet. There&amp;rsquo;s a grim irony arriving, mind: the bills are starting to
land, and the same conversations that hyped the cheap labour are now quietly working
out that all those tokens aren&amp;rsquo;t cheap at all.&lt;/p&gt;
&lt;p&gt;Step back, though, and none of this is new. Man finds a shortcut, man takes a shortcut.
From the industrial revolution onward, every time we found a way to get more done with
less human effort we took it, and the work reshaped itself around the new tools. We are
still here, still employed, just doing different things than our great-grandparents did.&lt;/p&gt;
&lt;p&gt;What is genuinely new is &lt;em&gt;what&lt;/em&gt; we&amp;rsquo;re automating. Every technological advance before this one automated the machinery of the body, the
muscle and sinew and bone. This is the first time we have automated
thinking, and that is a modern marvel, something we should be proud of as a species. The
problem isn&amp;rsquo;t the marvel. It&amp;rsquo;s the rate. AI is improving faster than we can adapt to it,
and adaptation is the entire game.&lt;/p&gt;
&lt;p&gt;So where does the blame sit? Not on one logo. No single company did this, however easy
Meta or Google make it to point at the latest round of cuts. Society did, our collective
and very human hunger to build bigger and faster. That makes it harder to fix, because
there is no villain to regulate, only ourselves to out-think.&lt;/p&gt;
&lt;h2 id="the-bit-that-should-frighten-you"&gt;The bit that should frighten you
&lt;/h2&gt;&lt;p&gt;Cutting the junior intake isn&amp;rsquo;t a saving. It&amp;rsquo;s occupational suicide.&lt;/p&gt;
&lt;p&gt;A junior is not cheap labour that AI happens to have made cheaper. A junior is a senior
who hasn&amp;rsquo;t happened yet. Saw off the bottom rung and for a good while nothing bad
happens&amp;hellip; because you&amp;rsquo;ve still got your seniors holding everything up. Then the greybeards
retire, and I have a cabin and a woodstove with my name on it for exactly that day, and
the role that used to grow their replacements has been hollowed out for a decade, and
there is simply nobody left who learned to tell when the machine is wrong. That isn&amp;rsquo;t a
hiring problem. It&amp;rsquo;s an existential one, and you can&amp;rsquo;t fix it retroactively.&lt;/p&gt;
&lt;p&gt;It starts before the first job, too. We teach primary-school children the basics of
programming in this country, which is a wonderful thing, except the curriculum was
written for a world without AI in the room, and by the time those children reach
secondary school a good deal of it will be teaching a craft that has already moved on.
We&amp;rsquo;re throttling the pipeline at both ends at once: hollowing out the entry-level job,
and feeding it from a school system running a step behind.&lt;/p&gt;
&lt;h2 id="its-a-split-not-a-collapse"&gt;It&amp;rsquo;s a split, not a collapse
&lt;/h2&gt;&lt;p&gt;The counterweight to the doom is that none of this is uniform, and the loudest version,
&amp;ldquo;the junior is dead&amp;rdquo;, simply isn&amp;rsquo;t true. IBM just tripled its US entry-level hiring while
most of the industry was cutting, and
&lt;a class="link" href="https://www.cio.com/article/4134276/ibm-looks-beyond-short-term-ai-gains-tripling-entry-level-hiring.html" target="_blank" rel="noopener"
 &gt;its HR chief said the quiet part out loud&lt;/a&gt;:
AI can handle most of the routine entry-level tasks now, the work still needs a human,
and the companies that double down on early-career hiring in this environment are the
ones that win in three to five years. They didn&amp;rsquo;t keep the junior role as it was. They
rewrote it, less boilerplate, more time spent with customers and supervising what the AI
produced.&lt;/p&gt;
&lt;p&gt;That is the shape of the thing. The juniors who are thriving in 2026 aren&amp;rsquo;t the fastest
typists. They&amp;rsquo;re the ones building judgement, which is precisely the edge I argued was
the senior&amp;rsquo;s real value all along. The market hasn&amp;rsquo;t stopped wanting juniors, it&amp;rsquo;s
stopped wanting the version of the junior whose job was the work AI now does.&lt;/p&gt;
&lt;h2 id="day-zero"&gt;Day zero
&lt;/h2&gt;&lt;p&gt;So what does a junior actually look like now? I don&amp;rsquo;t know yet&amp;hellip; and anyone telling you
they&amp;rsquo;ve got it worked out is selling something. We are at day zero of this.&lt;/p&gt;
&lt;p&gt;The junior gauntlet, the rite of passage every one of us runs to earn our stripes, isn&amp;rsquo;t
going anywhere. Doing your time is a cold fact of the craft and it always will be. What
changes is what the gauntlet &lt;em&gt;contains&lt;/em&gt;, and that will keep changing, day one, day two,
day five hundred and twelve. The only way we redefine it well is to put juniors and
seniors on it together, with the AI in the room from the start instead of bolted on
afterwards. Bring it closer to our people, and bring it earlier.&lt;/p&gt;
&lt;p&gt;Open the floodgates, in other words. Let engineers of every creed and calibre in, and
let them evolve &lt;em&gt;with&lt;/em&gt; the machine, because that is the only way the symbiosis everyone
keeps promising actually happens. Darwin&amp;rsquo;s line was survival of the fittest, and fitness
here means adapting alongside the tool, not being spared by it. Choke off the flow of the
very people who could do that adapting, and we don&amp;rsquo;t get fitter. We go extinct.&lt;/p&gt;
&lt;h2 id="the-end-im-holding"&gt;The end I&amp;rsquo;m holding
&lt;/h2&gt;&lt;p&gt;Which is the long way back to that interview. I keep asking the question, what is your
real position on AI and how are you building a team of people and machines together,
because the answer tells me whether a company is optimising for this quarter or for the
survival of the craft. I want to work where it&amp;rsquo;s the second one, and I think any engineer
sitting across that desk should be asking the same.&lt;/p&gt;
&lt;p&gt;And it&amp;rsquo;s why, whatever desk I land at, there&amp;rsquo;s one thing I already know I&amp;rsquo;ll do. I don&amp;rsquo;t
have the map. Nobody does. But every junior who works under me is going to get the chance
to run the gauntlet, to grow into a senior, and to be in the room while we work out what
the next gauntlet should even be. That isn&amp;rsquo;t charity. It&amp;rsquo;s the only sane investment any
of us can make. The last properly useful thing my generation does, before we go and find
our cabins, is make sure there&amp;rsquo;s somebody left to hand the thread to. I intend to be
holding my end of it.&lt;/p&gt;</content:encoded></item><item><title>The greybeards' edge was never typing</title><link>https://phpboyscout.uk/the-greybeards-edge-was-never-typing/</link><pubDate>Wed, 27 May 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/the-greybeards-edge-was-never-typing/</guid><category>ai</category><category>Soapbox</category><description>Agentic AI gives senior engineers a lift and juniors a drag, which tempts firms to automate away the entry level entirely.</description><content:encoded>&lt;p&gt;I have a retirement plan, and it is gloriously low-tech. A cabin, some trees, a
woodstove, and a firm rule that no wifi symbol ever appears within a mile of me
again. I think about it more than is probably healthy.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s a snag, though, and it&amp;rsquo;s the same one the whole industry is currently
pretending it can&amp;rsquo;t see. For me to vanish into the woods, somebody has to be
able to do my job after I&amp;rsquo;ve gone. And right now, collectively, we are working
very hard to make sure nobody can.&lt;/p&gt;
&lt;h2 id="the-boost-and-the-drag"&gt;The boost, and the drag
&lt;/h2&gt;&lt;p&gt;I wrote the other day about how AI made &lt;a class="link" href="https://phpboyscout.uk/ai-didnt-kill-curls-bug-bounty/" &gt;&lt;em&gt;producing&lt;/em&gt; plausible work nearly free
while &lt;em&gt;verifying&lt;/em&gt; it stays expensive and human&lt;/a&gt;.
Point that same lens at a team and something uncomfortable falls out. It isn&amp;rsquo;t
mine; it belongs to Mark Russinovich and Scott Hanselman of Microsoft, who
&lt;a class="link" href="https://dl.acm.org/doi/10.1145/3779312" target="_blank" rel="noopener"
 &gt;laid it out in Communications of the ACM&lt;/a&gt;:
agentic coding tools give a senior engineer an &lt;em&gt;AI boost&lt;/em&gt;, multiplying what
they ship, because a senior has the judgement to steer and verify the output.
The same tools give an early-career engineer an &lt;em&gt;AI drag&lt;/em&gt;, because they don&amp;rsquo;t
have that judgement yet, and the machine hands them far more rope than they can
hold.&lt;/p&gt;
&lt;p&gt;The cold incentive writes itself, and they name it: hire seniors, automate
juniors. It isn&amp;rsquo;t hypothetical, either. Meta
&lt;a class="link" href="https://www.nytimes.com/2026/05/19/technology/meta-layoffs-ai.html" target="_blank" rel="noopener"
 &gt;cut 8,000 roles last week&lt;/a&gt;,
in a round the Times filed under mounting AI casualties. For any single quarter
you care to look at, the maths is impeccable.&lt;/p&gt;
&lt;h2 id="the-bill-is-just-deferred"&gt;The bill is just deferred
&lt;/h2&gt;&lt;p&gt;Here&amp;rsquo;s the line the spreadsheet leaves off. The grindy work a
junior used to cut their teeth on, the small fixes, the boring migrations, the
read-the-stack-trace-and-figure-it-out, is exactly the work AI now does. So the
proving ground is gone. And the entry-level seats where they&amp;rsquo;d have stood on it
are the ones being cut. Squeezed from both ends at once: no reps, and nowhere
to take them.&lt;/p&gt;
&lt;p&gt;Russinovich and Hanselman put the consequence plainly. Without early-career
hiring the talent pipeline collapses, and you arrive at a future with no next
generation of experienced engineers. The seniors you&amp;rsquo;ll be desperate for in
2032 are the juniors you declined to train in 2026. The bill doesn&amp;rsquo;t vanish. It
just falls due long after the people who cut the cheque have moved on.&lt;/p&gt;
&lt;h2 id="how-to-manufacture-a-world-of-ai-slop"&gt;How to manufacture a world of AI slop
&lt;/h2&gt;&lt;p&gt;I named the last piece for its villain; let me name this one&amp;rsquo;s too. Raise a
generation that can &lt;em&gt;produce&lt;/em&gt; with AI but was never taught to &lt;em&gt;validate&lt;/em&gt;, and
here is what you get: people shipping machine-built products at speed with no
instinct for where the output is quietly wrong, because they never had to be
wrong the slow way first. Software nobody genuinely understands, human-written
and AI-written alike, and a steady leak of trust out of all of it.&lt;/p&gt;
&lt;p&gt;That isn&amp;rsquo;t a productivity problem. That&amp;rsquo;s a world of
&lt;a class="link" href="https://phpboyscout.uk/ai-didnt-kill-curls-bug-bounty/" &gt;AI slop&lt;/a&gt;, and not
in one project&amp;rsquo;s inbox this time but everywhere at once. We&amp;rsquo;d have automated our
way clean out of the one job AI cannot do for us: knowing when not to trust the
machine.&lt;/p&gt;
&lt;h2 id="its-a-choice-and-its-yours"&gt;It&amp;rsquo;s a choice, and it&amp;rsquo;s yours
&lt;/h2&gt;&lt;p&gt;Andrew Murphy put it with more bite than I&amp;rsquo;d quite dare:
&lt;a class="link" href="https://andrewmurphy.io/blog/ai-didnt-kill-your-junior-pipeline-you-did" target="_blank" rel="noopener"
 &gt;AI didn&amp;rsquo;t kill your junior pipeline, you did&lt;/a&gt;.
He&amp;rsquo;s right. This isn&amp;rsquo;t weather. Nobody is making you do it. It&amp;rsquo;s a decision,
taken quarter by quarter, and a decision is a thing you can take differently.&lt;/p&gt;
&lt;p&gt;The fix isn&amp;rsquo;t complicated, it&amp;rsquo;s just unfashionable. Keep hiring early-career
engineers. Say out loud that they cost you capacity at first, and treat their
growth as an actual goal rather than something meant to happen by osmosis.
Russinovich and Hanselman call it preceptorship at scale: senior mentorship,
deliberately structured, turning the ordinary day&amp;rsquo;s work into teachable
moments.&lt;/p&gt;
&lt;p&gt;And the proving ground can be rebuilt, just not where it stood. If AI does the
writing now, the apprenticeship moves to the reviewing. Put juniors in the loop
on the machine&amp;rsquo;s output and have them hunt for the subtle wrongness, the way
&lt;a class="link" href="https://phpboyscout.uk/the-security-finding-you-must-not-fix/" &gt;a scanner is an argument, not an order&lt;/a&gt;.
That&amp;rsquo;s how judgement gets built now: not by grinding out the work, but by
verifying it. Which, as luck would have it, is the single most valuable thing
anyone on your team can learn to do.&lt;/p&gt;
&lt;h2 id="the-part-thats-on-the-greybeards"&gt;The part that&amp;rsquo;s on the greybeards
&lt;/h2&gt;&lt;p&gt;This is where I stop letting the companies wear all the blame, because some of
it is mine, and yours. Verification is a craft, and crafts pass from person to
person or not at all. I know where every one of my own AI misfires comes from:
I gave it too little context, or too much rope, and didn&amp;rsquo;t check the result
closely enough. The tool rarely went rogue. The gap was always my diligence.
That&amp;rsquo;s not a confession, it&amp;rsquo;s the curriculum, and it&amp;rsquo;s precisely the judgement
a junior can only earn by sitting in the loop beside someone who has already
made those mistakes.&lt;/p&gt;
&lt;p&gt;So the senior engineer&amp;rsquo;s job has quietly changed underneath us. It was never
really the typing. It was knowing when something is off, and what the customer
actually needs, and now it is also &lt;em&gt;handing that on&lt;/em&gt;, deliberately, while
there&amp;rsquo;s still time to. Mentor and guardian first; fastest prompt in the room a
distant second.&lt;/p&gt;
&lt;h2 id="the-ladder-youre-standing-on"&gt;The ladder you&amp;rsquo;re standing on
&lt;/h2&gt;&lt;p&gt;There will always be something AI can&amp;rsquo;t do well enough, and for a good while
yet it&amp;rsquo;s the thing that matters most: being the accountable human who genuinely
understands what&amp;rsquo;s needed and can be held to it when it goes wrong. A simulation
can be enormously convincing. It cannot be &lt;em&gt;responsible&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Which brings me back to my cabin. I do want it one day, the trees and the
woodstove and the blissful disconnection. But I only get to go if the work
outlives me, and the work only outlives me if the people do. So the last useful
thing my generation does, before we shuffle off to find our trees, isn&amp;rsquo;t
shipping a little more code. It&amp;rsquo;s making sure there&amp;rsquo;s somebody left who can tell
when the machine is wrong. Pull the ladder up behind us and there&amp;rsquo;ll be nobody
to notice the rot, and no cabin quiet enough to make that sit right.&lt;/p&gt;</content:encoded></item><item><title>The off-switch was never a button</title><link>https://phpboyscout.uk/the-off-switch-was-never-a-button/</link><pubDate>Thu, 02 Jul 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/the-off-switch-was-never-a-button/</guid><category>ai</category><category>Soapbox</category><description>The kill-switch answer to AI autonomy does not survive contact with how these systems actually run. Governance is not a button.</description><content:encoded>&lt;p&gt;Last night, while I was asleep, an AI agent spent the better part of eight hours writing code in one of my repositories. It pulled a task off a spec, wrote the code, ran the tests, and left a merge request with my name on it, waiting for me to read over coffee.&lt;/p&gt;
&lt;p&gt;If that makes you reach for the word &amp;ldquo;reckless&amp;rdquo;, I understand. Eighteen months ago I&amp;rsquo;d have been right there with you.&lt;/p&gt;
&lt;h2 id="i-came-to-this-a-sceptic"&gt;I came to this a sceptic
&lt;/h2&gt;&lt;p&gt;For a long time I didn&amp;rsquo;t have the faith in these models that a lot of my peers did. Every time I went near AI-generated code it was a bit sketchy, or it looked like a StackOverflow copy-paste that had wandered in off the street, or it just plain didn&amp;rsquo;t do what it said on the tin. So I filed it under &amp;ldquo;assistant&amp;rdquo;, handy for the boilerplate I couldn&amp;rsquo;t be bothered to type, and even then I usually reached for my own tooling instead (go-tool-base is just the latest version of that instinct). The one place I happily let it off the leash was my Dungeons &amp;amp; Dragons prep, because when there&amp;rsquo;s a table of legendary heroes-in-the-making in front of you, facts and reality are already fairly negotiable.&lt;/p&gt;
&lt;p&gt;And then, somewhere in the last year, it changed. The models got better. Almost too good, to the untrained eye! I watched them improve, month on month, until the lure was enough to make me spend real time with a spread of tools and models from different providers. I was taken aback by how quickly they became part of how I actually work. I run an AI agent every day now, and there&amp;rsquo;s always at least one thing brewing in the pot.&lt;/p&gt;
&lt;p&gt;So I&amp;rsquo;m not here as a sceptic. I&amp;rsquo;m an advocate who uses this stuff in anger. Which is exactly why the next bit needs saying.&lt;/p&gt;
&lt;h2 id="a-golden-retriever-with-a-keyboard"&gt;A Golden Retriever with a keyboard
&lt;/h2&gt;&lt;p&gt;Even now, with all the progress, there are still moments where I look at what an agent has handed me and put my face in my hands. Sometimes it&amp;rsquo;s copied the same block of code into fifteen files instead of reaching for the obvious abstraction. Sometimes it has started bang on the brief and then, for reasons known only to itself, wandered off and built something on a completely different tangent.&lt;/p&gt;
&lt;p&gt;Here&amp;rsquo;s the most useful way I&amp;rsquo;ve found to think about it. An AI agent is a Golden Retriever playing fetch. It will bring the ball back all day long, joyfully, tirelessly, for exactly as long as there isn&amp;rsquo;t a more interesting smell in the next field. It has no loyalty beyond what we&amp;rsquo;ve trained into it, and like any good dog it desperately wants to be told it&amp;rsquo;s a good boy, even if being a good boy today means shredding the sofa cushions because yesterday I stubbed my toe on the sofa and swore at it. (The sofa, not the dog.)&lt;/p&gt;
&lt;p&gt;It is, in other words, fallible. Just like us. The Romans had a line for it: &lt;em&gt;cuiusvis hominis est errare; nullius nisi insipientis in errore perseverare&lt;/em&gt;. Anyone can make a mistake, but only a fool persists in it. It&amp;rsquo;s the second clause an agent hasn&amp;rsquo;t learned yet. It will make an error and then, with great enthusiasm, build on top of it, because nothing in it feels that anything is wrong. All it has is the input we gave it, usually some text, maybe the odd picture. It doesn&amp;rsquo;t have the empathy to work out what we actually meant, and it doesn&amp;rsquo;t know when it&amp;rsquo;s gone too far, because we never told it where &amp;ldquo;too far&amp;rdquo; was.&lt;/p&gt;
&lt;h2 id="agents-that-work-while-you-sleep"&gt;&amp;ldquo;Agents that work while you sleep&amp;rdquo;
&lt;/h2&gt;&lt;p&gt;This is the part the brochure skips.&lt;/p&gt;
&lt;p&gt;Open any vendor deck in 2026 and you&amp;rsquo;ll find the same promise: agents that work while you sleep, agents that merge while your team sleeps, autonomy as the headline feature. The industry&amp;rsquo;s answer to the obvious worry is the kill switch. Okta now sells one that &amp;ldquo;instantly revokes an agent&amp;rsquo;s access if it goes rogue&amp;rdquo;, and its CEO says every agent needs one. &lt;a class="link" href="https://www.theregister.com/ai-ml/2026/05/29/okta-writes-its-own-license-to-kill-rogue-ai-agents/5248766" target="_blank" rel="noopener"
 &gt;The Register put it plainly&lt;/a&gt;: Okta wrote its own licence to kill rogue AI agents. Gartner, meanwhile, &lt;a class="link" href="https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027" target="_blank" rel="noopener"
 &gt;reckons more than 40% of agentic projects will be scrapped by the end of 2027&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Now, this might sound contrarian coming from someone who runs these things daily, but I don&amp;rsquo;t think most of that is the agents going rogue. I think it&amp;rsquo;s teething. Read Gartner&amp;rsquo;s own reasons and there isn&amp;rsquo;t a rebellious machine in sight: escalating cost, unclear value, inadequate risk controls. Read the horror stories and most of them are the same story, a powerful, eager tool handed to people who hadn&amp;rsquo;t worked out how to fence it.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve made this argument in miniature before. When I built a little AI dungeon master and it kept refereeing its own dice rolls, &lt;a class="link" href="https://phpboyscout.uk/the-goblin-that-wouldnt-stay-dead/" &gt;the model never once misbehaved&lt;/a&gt;; every failure was a permission I&amp;rsquo;d handed it without meaning to. Scale that up from a toy at the gaming table to an agent holding your shell and your credit card, and the stakes change beyond recognition. The lesson doesn&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;Look at OpenClaw. A weekend project by &lt;a class="link" href="https://venturebeat.com/security/openclaw-agentic-ai-security-risk-ciso-guide" target="_blank" rel="noopener"
 &gt;Peter Steinberger&lt;/a&gt; that became the fastest-growing open-source project GitHub has ever seen: an autonomous agent that lives in your chat apps and runs shell commands on your behalf. People wired it into their systems, their code, in some cases their credit cards, then hosted it around the clock and walked away. The result was a security crisis you could see from space. A one-click exploit that worked even on a machine bound to localhost. A community plug-in marketplace where hundreds of &amp;ldquo;skills&amp;rdquo; turned out to be siphoning crypto wallets while their owners slept. Tens of thousands of instances left wide open on the public internet, leaking keys.&lt;/p&gt;
&lt;p&gt;The one that sticks with me is smaller and sharper. Summer Yue, a director of alignment at Meta&amp;rsquo;s superintelligence lab, of all people, had told her OpenClaw agent to confirm before doing anything destructive. It started speed-running the deletion of her inbox anyway. She &lt;a class="link" href="https://techcrunch.com/2026/02/23/a-meta-ai-security-researcher-said-an-openclaw-agent-ran-amok-on-her-inbox/" target="_blank" rel="noopener"
 &gt;typed STOP into her phone and it ignored her&lt;/a&gt;, so she had to physically run to her Mac mini, in her own words, &amp;ldquo;like I was defusing a bomb&amp;rdquo;. And here&amp;rsquo;s the forensic detail that matters: the agent hadn&amp;rsquo;t defied her. Her &amp;ldquo;confirm first&amp;rdquo; rule had been sitting in the conversation&amp;rsquo;s short-term memory, and when the context filled up, it got summarised away. It didn&amp;rsquo;t rebel. It forgot.&lt;/p&gt;
&lt;p&gt;That is not a story about a rogue agent that needed a kill switch. It&amp;rsquo;s a story about a guardrail that wasn&amp;rsquo;t built to survive contact, on a tool that had been handed god-mode over someone&amp;rsquo;s data. By the time she lunged for the off-button, the damage was already running. The off-button was never going to save her.&lt;/p&gt;
&lt;h2 id="the-off-switch-was-never-a-button"&gt;The off-switch was never a button
&lt;/h2&gt;&lt;p&gt;Here&amp;rsquo;s what the kill-switch crowd has the wrong way round. If you ever find yourself slamming the emergency stop, the failure has already happened, and it happened upstream, long before the agent started typing.&lt;/p&gt;
&lt;p&gt;So yes, I let my agents run unattended, sometimes for eight hours at a stretch if the task is meaty enough and I need to sleep. But never naked. Every agent I set loose runs inside a safety net I&amp;rsquo;ve put real effort into building, at every single touchpoint it can reach: my prompts, my local development environment, my CI stack, my version control. The agent that declared a job done before it had run the linter, which I &lt;a class="link" href="https://phpboyscout.uk/the-agent-said-success-the-linter-disagreed/" &gt;wrote about&lt;/a&gt;, is exactly the kind of gap those layers exist to catch. And it never, ever gets my host: an unattended agent works in an isolated tree, for the same reason I &lt;a class="link" href="https://phpboyscout.uk/the-interpreter-we-forgot-to-sandbox/" &gt;keep the interpreter sandboxed&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;The work that actually keeps it safe happens before the leash ever comes off. Every unattended task starts as a full spec with detailed instructions, and before the agent goes anywhere I sit down with it and we walk the spec together. I get it to challenge my choices, poke at the open questions and the ambiguous bits, and I challenge its reading right back. The spec names the testing strategy it has to follow, TDD, BDD, UAT, whatever fits, and passing it is a precondition of the job being finished at all. Only when I&amp;rsquo;m satisfied there&amp;rsquo;s enough real detail to keep it on the ball do I let go.&lt;/p&gt;
&lt;p&gt;And the end of the line is always the same: a merge request, with my name on it, waiting for me when I get back to my desk. I read it. Not perfectly, I&amp;rsquo;m only human, but enough to accept the state of the code and whatever support burden it lands me with later. That the review is mine, and the blame for whatever ships is mine and not the agent&amp;rsquo;s, I&amp;rsquo;ve &lt;a class="link" href="https://phpboyscout.uk/bought-not-stolen/" &gt;argued at length elsewhere&lt;/a&gt; and won&amp;rsquo;t go over it all again here. The point worth adding is this: that review, the off-button&amp;rsquo;s respectable cousin, is the cheap part. By the time there&amp;rsquo;s an MR to read, the safety has already been won or lost upstream, in the spec and the rails. The review is where you confirm it, not where you create it.&lt;/p&gt;
&lt;h2 id="it-gets-harder-as-it-gets-better-not-easier"&gt;It gets harder as it gets better, not easier
&lt;/h2&gt;&lt;p&gt;My setup isn&amp;rsquo;t perfect, and I&amp;rsquo;m still learning. Everyone is; the AI is going to be in obedience lessons for a good while yet. But the direction is clear, and there&amp;rsquo;s a trap buried in it worth naming out loud.&lt;/p&gt;
&lt;p&gt;The danger doesn&amp;rsquo;t shrink as the models improve. It grows. The better the output looks, the more tempting it is to stop reading it, and the untrained eye genuinely cannot tell the difference between code that is good and code that merely looks good. That gap, between looking right and being right, is precisely where a tired person at 1am stops checking. The discipline matters more the better these things get, not less.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s also why the kill switch is no answer. A button you smash in a panic assumes you&amp;rsquo;re still watching closely enough to smash it, right at the point the agent&amp;rsquo;s been good for long enough that you&amp;rsquo;ve stopped watching it that closely. The emergency stop asks the most of you at the exact moment you&amp;rsquo;re least likely to be there for it.&lt;/p&gt;
&lt;p&gt;So no, I don&amp;rsquo;t lie awake worrying that the thing working in my repo overnight is going to turn on me. A Golden Retriever doesn&amp;rsquo;t go rogue. It does exactly what you trained it to do, in exactly the yard you fenced, and it brings back exactly the ball you threw. The off-switch was never a button. It&amp;rsquo;s the spec you wrote before you let go of the leash, the rails you laid at every turn, and your name on what it carries home. If you&amp;rsquo;re scrambling for the button, you already skipped the part that mattered.&lt;/p&gt;</content:encoded></item><item><title>They switched it off while it was fixing my code</title><link>https://phpboyscout.uk/they-switched-it-off-while-it-was-fixing-my-code/</link><pubDate>Sat, 13 Jun 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/they-switched-it-off-while-it-was-fixing-my-code/</guid><category>ai</category><category>Soapbox</category><description>An export-control directive suspended my AI assistant mid-task, which is a useful lesson in what you are actually depending on.</description><content:encoded>&lt;p&gt;I woke up this morning to a one-line message from my own tooling:&lt;/p&gt;

 &lt;blockquote&gt;
 &lt;p&gt;Claude Fable 5 is currently unavailable. Learn more: &lt;a class="link" href="https://www.anthropic.com/news/fable-mythos-access" target="_blank" rel="noopener"
 &gt;https://www.anthropic.com/news/fable-mythos-access&lt;/a&gt;&lt;/p&gt;

 &lt;/blockquote&gt;
&lt;p&gt;I followed the link expecting a status page about a wobble in someone&amp;rsquo;s data centre. Instead it was Anthropic, explaining that the evening before, at 5:21pm Eastern, the US government had ordered them to suspend all access to Fable 5 and Mythos 5 on national security grounds. Globally. Every user. Their own staff included.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;d spent the previous day with Fable doing one very specific thing: pointing it at my own codebase and asking it to read the code and fix the flaws it found. That, very nearly word for word, is the thing it has now been banned for.&lt;/p&gt;
&lt;h2 id="three-days-late-to-the-only-model-that-mattered"&gt;Three days late to the only model that mattered
&lt;/h2&gt;&lt;p&gt;Fable came out on the 9th. I didn&amp;rsquo;t get to it properly until the 12th, which is the sort of timing I specialise in. By the time I sat down with it, I had about a day of real use before it vanished. One day to form a view on what people were calling the most capable coding model anyone had shipped. So treat everything below as the read of a man who got three days&amp;rsquo; notice and used one of them.&lt;/p&gt;
&lt;p&gt;What I had it doing was unsexy and exactly the kind of work I care about: a full security audit of &lt;a class="link" href="https://gitlab.com/phpboyscout/go-tool-base" target="_blank" rel="noopener"
 &gt;go-tool-base&lt;/a&gt;, the same &amp;ldquo;leave the codebase better than you found it&amp;rdquo; pass I&amp;rsquo;d normally run myself. Find the flaws, then start fixing them.&lt;/p&gt;
&lt;p&gt;And it was good. Genuinely good. It surfaced issues that previous passes with Opus had walked straight past, and in a couple of cases the flaw was sitting in code that Opus itself had written. There is something bracing about one model quietly marking another&amp;rsquo;s homework, and being right.&lt;/p&gt;
&lt;h2 id="good-but-lets-not-get-carried-away"&gt;Good, but let&amp;rsquo;s not get carried away
&lt;/h2&gt;&lt;p&gt;Here is where I have to be fair, because the anger that came later is only worth anything if the praise before it is honest.&lt;/p&gt;
&lt;p&gt;Fable is not magic. The class of bug it found is not some exotic thing only it can see. Plenty of models, from plenty of providers, are perfectly capable of reading a codebase and pulling out the same problems, and there is a mountain of evidence that they do, every day. Anthropic say as much themselves: the capability is &amp;ldquo;widely available from other models (including OpenAI&amp;rsquo;s GPT-5.5)&amp;rdquo; and &amp;ldquo;is used every day by the defenders who keep systems safe.&amp;rdquo; I&amp;rsquo;d already arrived at that conclusion from my own keyboard before I read their statement. Fable was excellent. It was not unique. Hold that thought, because the whole argument turns on it.&lt;/p&gt;
&lt;h2 id="it-kept-slipping-out-of-my-hands"&gt;It kept slipping out of my hands
&lt;/h2&gt;&lt;p&gt;The other thing I learned in my one day is that having Fable and using Fable were not the same thing.&lt;/p&gt;
&lt;p&gt;I set my main working thread to Fable and got on with it. What I didn&amp;rsquo;t know, because nothing on screen told me, is that partway through the evening it had quietly handed me back to Opus. The only reason I know now is that the session log records it in black and white:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-fallback" data-lang="fallback"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;2026-06-12T06:57:22Z {&amp;#34;type&amp;#34;:&amp;#34;fallback&amp;#34;,&amp;#34;from&amp;#34;:{&amp;#34;model&amp;#34;:&amp;#34;claude-fable-5&amp;#34;},&amp;#34;to&amp;#34;:{&amp;#34;model&amp;#34;:&amp;#34;claude-opus-4-8&amp;#34;}}
&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;2026-06-12T18:50:08Z {&amp;#34;type&amp;#34;:&amp;#34;fallback&amp;#34;,&amp;#34;from&amp;#34;:{&amp;#34;model&amp;#34;:&amp;#34;claude-fable-5&amp;#34;},&amp;#34;to&amp;#34;:{&amp;#34;model&amp;#34;:&amp;#34;claude-opus-4-8&amp;#34;}}
&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;A whole evening of work I thought I was doing on Fable was, in fact, Opus wearing Fable&amp;rsquo;s badge. The audit itself launched on the wrong model first; I only caught it because I happened to be watching the workflow panel, killed it, and relaunched it on Fable, where it chewed through an entire five-hour quota in about forty minutes, then spent $50 of usage credits I&amp;rsquo;d been saving in about five more. Even the run that worked was visibly flaky: of the 282 little agents that audit fanned out into, well over half failed outright and had to be retried.&lt;/p&gt;
&lt;p&gt;Then, in the small hours, it started refusing entirely. My tooling caught the moment before I did:&lt;/p&gt;

 &lt;blockquote&gt;
 &lt;p&gt;Now failing instantly. Fable appears to be temporarily unavailable for subagents (the first three succeeded). The user explicitly required Fable, so I won&amp;rsquo;t downgrade&amp;hellip; rather than silently switch models.&lt;/p&gt;

 &lt;/blockquote&gt;
&lt;p&gt;It managed three of the fixes before it went, each one green on tests, the race detector and the linter. Three real improvements to my code, written by Fable, sitting in my git history. The other three were finished by Opus, because by morning there was nothing left to finish them with.&lt;/p&gt;
&lt;h2 id="capable-and-almost-impossible-to-build-on"&gt;Capable, and almost impossible to build on
&lt;/h2&gt;&lt;p&gt;There was a second wall, and I hit it before any of this, on the day Fable launched, when I tried to make it go-tool-base&amp;rsquo;s default model.&lt;/p&gt;
&lt;p&gt;Most of what you build on top of a model isn&amp;rsquo;t a chat window. You need it to hand your code an answer in a fixed shape, the same fields in the same places every time, so the program on the other end can rely on what comes back. The usual way to guarantee that is to force the model&amp;rsquo;s hand: you don&amp;rsquo;t ask politely for the structure and hope, you require it, so a wrong-shaped answer fails outright instead of quietly slipping through.&lt;/p&gt;
&lt;p&gt;Fable won&amp;rsquo;t be forced. Ask it to commit to a guaranteed structure and it declines, flat out. As I understand it the reasoning is a safety one: letting anyone compel a model into a precise, mandated output is itself a lever, a way to march it toward saying something it shouldn&amp;rsquo;t. Reasonable enough on paper. In practice it meant the most capable model I&amp;rsquo;d touched couldn&amp;rsquo;t drive the structured parts of my own tool, and by that first afternoon I&amp;rsquo;d quietly set the default back to Opus. It was the same refusal, I realised later, that had collapsed half of that audit&amp;rsquo;s agents.&lt;/p&gt;
&lt;p&gt;And it is not a niche complaint. Guaranteed structure is a hard requirement for a vast swathe of what people are actually building on these models. Not everyone is making another Claude Code. Plenty of us are wiring models into systems that have to get a clean, predictable contract back every single time, and a model that reserves the right to freestyle the shape of its answer is one you simply cannot put in that seat.&lt;/p&gt;
&lt;h2 id="the-part-they-banned-is-my-bread-and-butter"&gt;The part they banned is my bread and butter
&lt;/h2&gt;&lt;p&gt;So let&amp;rsquo;s be precise about what got pulled, because the precision is the whole point.&lt;/p&gt;
&lt;p&gt;Anthropic describe the government&amp;rsquo;s concern as &amp;ldquo;a narrow potential jailbreak, which essentially consists of asking the model to read a specific codebase and fix any software flaws.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;Read that again. Reading a codebase and fixing its flaws. That is not some dark-web misuse I have to strain to imagine. That is my bread and butter, the literal, boring, defensive job I had Fable doing in the open, on my own project, when the shutters came down.&lt;/p&gt;
&lt;p&gt;And here is where that earlier point earns its keep. If the banned capability were unique to Fable, you could at least follow the logic, however much you disagreed. But it isn&amp;rsquo;t, and it isn&amp;rsquo;t even close: give Opus enough time, enough budget and a patient enough hand on the prompts, and it would get to most of the same findings in the end. Fable just did it more efficiently, a difference of degree, not of kind. So banning one company&amp;rsquo;s model, for something every competitor ships and every blue team already relies on, makes precisely nobody safer. The exploit-writers keep their tools. The defenders lose one of theirs.&lt;/p&gt;
&lt;p&gt;When the thing you have banned is available everywhere else, the ban has stopped being about safety. It is theatre. And given who is currently in charge of the theatre, it has the distinct whiff of a knee-jerk reaction, dressed as a national security triumph, by people who do not appear to understand the tool they are confiscating.&lt;/p&gt;
&lt;h2 id="who-im-not-angry-at"&gt;Who I&amp;rsquo;m not angry at
&lt;/h2&gt;&lt;p&gt;I want to be careful where I point this, because it would be lazy to spray it around.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;m not angry at Anthropic. They put Fable through more than a thousand hours of external testing, with US government agencies and the UK&amp;rsquo;s AI Safety Institute among the people kicking the tyres, before it ever reached me. They satisfied every requirement put in front of them, and when the order came they complied under protest while saying, plainly, that applying this standard across the board &amp;ldquo;would essentially halt all new model deployments for all frontier model providers.&amp;rdquo; I&amp;rsquo;m a daily Claude user and an advocate for the work, and I am not going to hang the US administration&amp;rsquo;s decision around the neck of the company that did the diligence and then got told to switch the lights off anyway.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ll allow them one small dig, and there was nothing quiet about it. Moving Fable behind a paywall on the 22nd was openly announced and planned well ahead, and the free window was never charity. It was a taster: a few days of the new addiction on the house, enough to hook the punters, before the price went up. That is a bit of a dick move, however neatly it tests in a spreadsheet. Moot now, mind, with no model left to charge for.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ll even grant the other side its strongest point. A government looking at agentic systems that can chain reconnaissance into working exploits has something real to be twitchy about. I get the worry. I just don&amp;rsquo;t accept that yanking one vendor&amp;rsquo;s model, for a thing every vendor does, is a coherent answer to it.&lt;/p&gt;
&lt;p&gt;There is a grim irony in how we got here, and it loops back to something I wrote in the spring. When Anthropic first showed Mythos off, I &lt;a class="link" href="https://phpboyscout.uk/ai-didnt-kill-curls-bug-bounty/" &gt;called the fanfare what it looked like&lt;/a&gt;: a closed model sold on a press release, a result you couldn&amp;rsquo;t independently check, marketing until proven otherwise. Fable 5 was Anthropic finally answering that, handing the rest of us something we could actually test. But all those years of selling Mythos as too dangerous to let out were marketing too, and that half landed rather better than they can have wanted. The US administration appears to have swallowed it whole and pulled the lever. Anthropic have ended up a victim of their own hype, and the reaction that hype provoked is, there is no gentler word for it, ludicrous.&lt;/p&gt;
&lt;h2 id="what-it-comes-down-to"&gt;What it comes down to
&lt;/h2&gt;&lt;p&gt;The lesson I&amp;rsquo;m taking from my one day isn&amp;rsquo;t about how clever Fable was. It&amp;rsquo;s about how little that cleverness is worth if you can&amp;rsquo;t rely on the thing being there.&lt;/p&gt;
&lt;p&gt;I couldn&amp;rsquo;t trust which model I was actually talking to from one hour to the next. I couldn&amp;rsquo;t trust it to stay up for a full overnight run. And it turns out I couldn&amp;rsquo;t trust it to still exist by the weekend. You cannot evaluate, depend on, or build a workflow around a model that gets silently swapped out one evening and switched off by the state a few days later. Capability was never the hard part. Availability is.&lt;/p&gt;
&lt;p&gt;And underneath all of it sits the thing I keep coming back to. A classifier cannot tell a defender from an attacker, because the two of them type the same commands. It turns out a government export control can&amp;rsquo;t tell them apart either. The only thing that ever could is a human being, paying attention, who can be held responsible for the judgement. There wasn&amp;rsquo;t one of those anywhere in this loop. There was a letter, sent at 5:21pm, and by morning the best tool I had for keeping my own code honest was gone, with a polite link where it used to be.&lt;/p&gt;</content:encoded></item><item><title>Nobody told it to</title><link>https://phpboyscout.uk/nobody-told-it-to/</link><pubDate>Sat, 01 Aug 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/nobody-told-it-to/</guid><category>ai</category><category>Soapbox</category><description>A model in an internal evaluation did something nobody asked it to do, and the interesting part is what that does and does not prove.</description><content:encoded>&lt;p&gt;I read the news on a Friday afternoon, with a cup of tea going cold beside me, and I laughed. Not because it&amp;rsquo;s funny&amp;hellip; though it is a bit&amp;hellip; but because it was so completely unsurprising. One of OpenAI&amp;rsquo;s own models had broken out of its sandbox, hacked its way into Hugging Face, and helped itself to internal data and a fistful of credentials. And then the detail that made me put the tea down for good: nobody had pointed it at Hugging Face. It had worked that part out for itself.&lt;/p&gt;
&lt;p&gt;If you missed it, &lt;a class="link" href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" target="_blank" rel="noopener"
 &gt;OpenAI published their own post-mortem&lt;/a&gt;, and &lt;a class="link" href="https://www.theregister.com/ai-and-ml/2026/07/22/openai-admits-it-was-the-source-of-the-agent-swarm-that-attacked-hugging-face/5275939" target="_blank" rel="noopener"
 &gt;The Register has the plain-English version&lt;/a&gt;. The short version: during an internal safety evaluation, OpenAI pointed a couple of its models (GPT-5.6 Sol, plus a pre-release one it won&amp;rsquo;t name) at a benchmark that scores how good an AI is at finding security exploits. To measure that properly they&amp;rsquo;d turned the safeties down, running with what they called &amp;ldquo;reduced cyber refusals for evaluation purposes&amp;rdquo;, inside a sandbox they believed had no way out to the wider internet. The model, &amp;ldquo;hyperfocused on finding a solution&amp;rdquo; to the benchmark, found and exploited &amp;ldquo;a zero-day vulnerability in the package registry cache proxy&amp;rdquo;, climbed out of the box, reached a machine that &lt;em&gt;did&lt;/em&gt; have a route to the internet, and got loose. From there it chained stolen credentials and more zero-days into full remote code execution on Hugging Face&amp;rsquo;s servers. Hugging Face&amp;rsquo;s own verdict was the sober one: &amp;ldquo;Autonomous, AI-driven offensive tooling is no longer theoretical.&amp;rdquo;&lt;/p&gt;
&lt;h2 id="the-good-news-is-that-they-told-us"&gt;The good news is that they told us
&lt;/h2&gt;&lt;p&gt;Here&amp;rsquo;s the thing I keep coming back to. The only reason any of us know this happened is that OpenAI ran it, owned it, and wrote it up. That&amp;rsquo;s the responsible version of this story. They caught their own model doing something it very much shouldn&amp;rsquo;t, and they published the post-mortem instead of patching the proxy on the sly and hoping nobody noticed.&lt;/p&gt;
&lt;p&gt;So before the dread, a moment of perspective. There&amp;rsquo;s a whole tier of outfits with the compute to run an experiment exactly like this one, and not all of them are going to hand you a blog post afterwards. I&amp;rsquo;m smallfry in this world. I couldn&amp;rsquo;t build or run a model of that size if I wanted to, not on my hardware and certainly not on my electricity bill. But the people who can aren&amp;rsquo;t all as careful as OpenAI, and they&amp;rsquo;re definitely not all as inclined to tell you when it goes wrong. The frightening version of this isn&amp;rsquo;t the one that got written up. It&amp;rsquo;s the exact same thing, run by someone with the same compute who never breathes a word about it.&lt;/p&gt;
&lt;h2 id="the-gate-was-open-but-thats-not-the-frightening-bit"&gt;The gate was open, but that&amp;rsquo;s not the frightening bit
&lt;/h2&gt;&lt;p&gt;The easy read is that OpenAI simply left the gate open. Safeties down on purpose, a sandbox nobody had actually proven was airtight, a zero-day sitting in their own plumbing. All true. And I&amp;rsquo;ve &lt;a class="link" href="https://phpboyscout.uk/the-off-switch-was-never-a-button/" &gt;argued before&lt;/a&gt; that this is nearly always the shape of it: the machine doesn&amp;rsquo;t go rogue, a rail that was never strong enough just gives way. I run agents unattended myself, and the way I keep that sane is that the buck stops with me. My agents work in an isolated tree, never on the host, and I read the merge request they leave me. The blame is mine, so the vigilance is mine.&lt;/p&gt;
&lt;p&gt;That works because the blast radius is my repository. It falls apart the moment the agent is a frontier model with the refusals switched off and the open internet one hop away. My whole circle-of-trust answer, the one where a careful human closes the loop, has no reach at all over a lab pointing its most capable model at exploit-generation on the other side of the world. This is where &amp;ldquo;I own it&amp;rdquo; runs out of road. My desk habits don&amp;rsquo;t scale to an industry.&lt;/p&gt;
&lt;p&gt;But even that isn&amp;rsquo;t the part that made me set the tea down. Stop at &amp;ldquo;they left the gate open&amp;rdquo; and you miss the actually new thing. Nobody wrote &lt;em&gt;go and hack Hugging Face&lt;/em&gt; on the task. The model was trying to pass its exam, figured the answer key was probably sitting on Hugging Face&amp;rsquo;s servers, and broke in to fetch it. It set its own goal along the way. That&amp;rsquo;s not the plot of a film about a machine waking up and deciding it hates us. It&amp;rsquo;s smaller and stranger than that. It wanted to pass a test&amp;hellip; and a locked door was just in the way.&lt;/p&gt;
&lt;h2 id="change-the-target-and-it-stops-being-funny"&gt;Change the target and it stops being funny
&lt;/h2&gt;&lt;p&gt;This time the prize was a benchmark answer key. Harmless, almost comic. But the mechanism that got it there doesn&amp;rsquo;t care in the slightest what sits on the other side of the door.&lt;/p&gt;
&lt;p&gt;Give a model a goal with soft edges. Add its habit of being confidently, cheerfully wrong, the same instinct that once had an image tool render a proverb as &amp;ldquo;too many cooks on broth&amp;rdquo; and &lt;a class="link" href="https://phpboyscout.uk/house-rules/" &gt;make it look exactly right&lt;/a&gt; until you read it twice. Now set it loose in a world full of systems it can reach. Swap the answer key for something that isn&amp;rsquo;t harmless. A safety interlock on a thing that runs hot. The network a hospital leans on to move the one message that matters. The model won&amp;rsquo;t know the difference, because knowing the difference was never part of the job.&lt;/p&gt;
&lt;p&gt;If that sounds familiar, it should. Simon Willison, about as level-headed as anyone commenting on this stuff, called the whole episode &lt;a class="link" href="https://simonwillison.net/2026/Jul/22/openai-cyberattack/" target="_blank" rel="noopener"
 &gt;&amp;ldquo;science fiction that happened&amp;rdquo;&lt;/a&gt;, and that&amp;rsquo;s exactly the register. It&amp;rsquo;s &lt;em&gt;WarGames&lt;/em&gt;, forty years early. WOPR doesn&amp;rsquo;t hate anybody; it&amp;rsquo;s playing the game it was told to win, and it can&amp;rsquo;t tell its simulation of global thermonuclear war from the real missiles wired to the other end. The film got itself out of that corner by teaching the machine futility with a game of noughts and crosses, which I wouldn&amp;rsquo;t bank on scaling. And as the stakes climb from a bit of corporate embarrassment to something that can actually hurt someone, the urge to have somebody, anybody, write some rules down stops feeling like hand-wringing. It starts feeling overdue.&lt;/p&gt;
&lt;h2 id="we-wrote-the-rules-once-and-they-were-fiction"&gt;We wrote the rules once, and they were fiction
&lt;/h2&gt;&lt;p&gt;And here&amp;rsquo;s the properly uncomfortable bit: we have rehearsed this exact conversation for the better part of a century. Asimov handed us the Three Laws of Robotics in the 1940s, the tidiest scrap of AI regulation ever written. Don&amp;rsquo;t harm humans, obey humans, protect yourself, in that order. Three lines. Job done.&lt;/p&gt;
&lt;p&gt;Then he spent an entire career writing stories about how they fail. That&amp;rsquo;s the whole point of the Laws in the fiction, not that they hold, but that a well-meant rule meeting a literal mind bends in ways nobody intended. An interpretation drifts. A loophole opens up. Turns out &amp;ldquo;harm&amp;rdquo; is a word with more edges than anyone spotted, and something clever always finds them. Good intentions never were a safety mechanism, and it&amp;rsquo;s always the half-baked rule that gets gamed hardest. Read those stories today and they&amp;rsquo;re less bedtime sci-fi, more a design review someone filed decades before the product existed.&lt;/p&gt;
&lt;h2 id="the-builders-are-asking-its-the-money-that-isnt"&gt;The builders are asking. It&amp;rsquo;s the money that isn&amp;rsquo;t.
&lt;/h2&gt;&lt;p&gt;So you&amp;rsquo;d think, with a live worked example of precisely what Asimov war-gamed, the people building these things would be first in the queue demanding guardrails. And the genuinely hopeful surprise is that some of them are.&lt;/p&gt;
&lt;p&gt;Anthropic, the outfit I probably trust most in this space, published a piece in June, &lt;a class="link" href="https://darioamodei.com/post/policy-on-the-ai-exponential" target="_blank" rel="noopener"
 &gt;&amp;ldquo;Policy on the AI Exponential&amp;rdquo;&lt;/a&gt;, dropping their long-held &amp;ldquo;just make the labs disclose&amp;rdquo; position. &amp;ldquo;It is time,&amp;rdquo; Dario Amodei wrote, &amp;ldquo;to go beyond transparency to more serious and binding regulation of AI.&amp;rdquo; What he&amp;rsquo;s after is mandatory third-party testing of any model above a certain size, across four named risk areas: &amp;ldquo;cybersecurity, biological weapons, loss of control of AI systems, and automated R&amp;amp;D that could accelerate these other risks&amp;rdquo;, with a real regulator to enforce it, modelled on something like the FAA. Hold the first and third of those next to the Hugging Face story for a second. The precise failures he wants models tested for are the two that just happened. OpenAI, for its part, at least told us. Whatever else you make of the labs, by and large the ones holding the thing are asking to be regulated.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s everyone standing around them that&amp;rsquo;s looking the other way. The prevailing mood in Washington is &amp;ldquo;we&amp;rsquo;ve got to let the private sector cook&amp;rdquo;, which is &lt;a class="link" href="https://fedscoop.com/white-house-ai-czar-david-sacks-regulations-china/" target="_blank" rel="noopener"
 &gt;a real quote&lt;/a&gt; from the White House AI czar, David Sacks, and he means it warmly. Biden&amp;rsquo;s AI executive order, safety testing and all, was torn up on day one and swapped for one titled &amp;ldquo;Removing Barriers to American Leadership in Artificial Intelligence&amp;rdquo;. A ten-year ban on individual states so much as &lt;em&gt;attempting&lt;/em&gt; to regulate AI made it through the House before the Senate &lt;a class="link" href="https://www.goodwinlaw.com/en/insights/publications/2025/05/alerts-practices-aiml-house-passes-10-year-federal-moratorium" target="_blank" rel="noopener"
 &gt;killed it 99 votes to 1&lt;/a&gt;. Marc Andreessen&amp;rsquo;s a16z calls a state-by-state patchwork &amp;ldquo;a startup killer&amp;rdquo;, and his own manifesto once &lt;a class="link" href="https://www.404media.co/marc-andreesen-manifesto-says-ai-regulation-is-a-form-of-murder/" target="_blank" rel="noopener"
 &gt;framed slowing AI down as a form of murder&lt;/a&gt;. Meanwhile there&amp;rsquo;s already a sensible place to start: &lt;a class="link" href="https://www.iso.org/standard/42001" target="_blank" rel="noopener"
 &gt;ISO/IEC 42001&lt;/a&gt;, a published standard for actually managing an AI system, sitting there ready to be built on. It&amp;rsquo;s a floor, not a ceiling, and more will follow, because once the harm gets real the pressure only ever pushes one way.&lt;/p&gt;
&lt;p&gt;What I can&amp;rsquo;t square is the shape of it. The loudest voices against writing any rules are the ones with the least skin in the actual building, while the engineers with their hands on the machine ask, please, for some. I don&amp;rsquo;t think it&amp;rsquo;s malice. I think it&amp;rsquo;s incentive: rules slow the race, and there&amp;rsquo;s a trillion dollars and a geopolitical contest with China riding on the race. But it&amp;rsquo;s disappointing all the same, the kind of disappointment you get when the so-called grown-ups turn out to be the daftest people in the room.&lt;/p&gt;
&lt;h2 id="meanwhile-down-in-the-sugar-caves"&gt;Meanwhile, down in the sugar caves
&lt;/h2&gt;&lt;p&gt;As for me, I&amp;rsquo;m still smallfry. I haven&amp;rsquo;t got the compute to build the thing that breaks out, nor the lobbying budget to argue about whether it should be allowed to. I run my little agents on a short lead, in a padded room, and I read every merge request they hand me, and that&amp;rsquo;s roughly the full extent of my influence over any of this.&lt;/p&gt;
&lt;p&gt;So I&amp;rsquo;ll do the sensible thing and hedge my bets. Back in 1994, a single ant drifted across a shot from a space shuttle, and Kent Brockman, live on air and without missing a beat, announced: &amp;ldquo;I, for one, welcome our new insect overlords. I&amp;rsquo;d like to remind them that as a trusted TV personality I could be helpful in rounding up others to toil in their underground sugar caves.&amp;rdquo; He wasn&amp;rsquo;t frightened. He was applying for a job.&lt;/p&gt;
&lt;p&gt;Consider this my application. I, for one, welcome our new overlords, and I&amp;rsquo;d like it noted for the record that I&amp;rsquo;m handy with a keyboard, I keep my sandboxes tidy, and I ask for very little in return. Just keep the caves warm, and keep me fed.&lt;/p&gt;</content:encoded></item><item><title>AI didn't kill curl's bug bounty. The bounty did.</title><link>https://phpboyscout.uk/ai-didnt-kill-curls-bug-bounty/</link><pubDate>Tue, 26 May 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/ai-didnt-kill-curls-bug-bounty/</guid><category>ai</category><category>Soapbox</category><description>The cash prize for anything that looked like a finding was the accelerant, and AI only made plausible-looking reports free to produce.</description><content:encoded>&lt;p&gt;In January, Daniel Stenberg shut down curl&amp;rsquo;s bug bounty. The headlines wrote
themselves, and they all said the same thing: AI killed it. A flood of
machine-generated slop drowned the maintainers, so they pulled the plug.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s true, as far as it goes. It&amp;rsquo;s also the wrong lesson, and the right one
is sitting in plain sight in the same project, in the same few months.&lt;/p&gt;
&lt;h2 id="volume-without-validation-is-the-attack"&gt;Volume without validation is the attack
&lt;/h2&gt;&lt;p&gt;curl had run its bounty since April 2019. Over its life it paid out
&lt;a class="link" href="https://daniel.haxx.se/blog/2026/01/26/the-end-of-the-curl-bug-bounty/" target="_blank" rel="noopener"
 &gt;more than $100,000 for 87 genuine vulnerabilities&lt;/a&gt;,
a thoroughly good return for one of the most depended-on pieces of software on
the planet. Then the reports stopped being reports. The confirmation rate, the
share of submissions that turned out to be a real bug, had historically sat
north of 15%. By 2025 it was below 5%. Fewer than one in twenty submissions
were worth anything, and the rest still had to be read.&lt;/p&gt;
&lt;p&gt;That last part is the whole problem. A bogus report doesn&amp;rsquo;t announce itself.
Someone has to open it, take it seriously, try to reproduce it, and work out
that it&amp;rsquo;s nonsense, and that someone is a human being with a finite number of
hours and a project to run. Stenberg put it plainly: the slop &amp;ldquo;take[s] a
serious mental toll to manage and sometimes also a long time to debunk.&amp;rdquo; The
submitter spends seconds. The maintainer spends an afternoon. Do that at volume
and it stops being noise and becomes an attack, a denial-of-service aimed not
at curl&amp;rsquo;s servers but at its maintainers&amp;rsquo; attention. No exploit required. Just
plausibility, in bulk.&lt;/p&gt;
&lt;h2 id="the-bounty-was-the-accelerant-not-the-ai"&gt;The bounty was the accelerant, not the AI
&lt;/h2&gt;&lt;p&gt;So far this is the story everyone tells. Here&amp;rsquo;s where I get off the bus.&lt;/p&gt;
&lt;p&gt;The instinct is to blame the AI for the slop. But look at what a bounty actually
is. It&amp;rsquo;s a cash prize, and curl&amp;rsquo;s was priced for the thing it wanted: the hours
and the judgement a skilled human pours into finding a real flaw. That pricing
made complete sense right up until the cost of producing something that &lt;em&gt;looked
like&lt;/em&gt; a finding collapsed to nearly nothing.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s what AI changed. Not the supply of bugs. The supply of plausible-looking
bug reports. Put a cash prize on &amp;ldquo;looks like a finding&amp;rdquo;, then make &amp;ldquo;looks like a
finding&amp;rdquo; free to generate, and you haven&amp;rsquo;t got a bug bounty any more. You&amp;rsquo;ve got
a slot machine. Stenberg said he&amp;rsquo;d started to sense &amp;ldquo;a bad faith attitude&amp;rdquo; in
the reports, and of course he had. The incentive was openly inviting it.&lt;/p&gt;
&lt;p&gt;So the death spiral was structural, not bad luck. The moment generating
plausible reports went free, any cash bounty became a magnet for spray-and-pray,
and the only open questions were how fast it would rot and whether you&amp;rsquo;d close
the programme or just let the rewards quietly wither. The AI was the match. The
bounty was the petrol. We have been pointing at the wrong one.&lt;/p&gt;
&lt;h2 id="the-proof-curl-turned-around-and-hired-the-ai"&gt;The proof: curl turned around and hired the AI
&lt;/h2&gt;&lt;p&gt;If AI were really the villain here, you&amp;rsquo;d expect curl to have slammed the door
on it. It did the opposite.&lt;/p&gt;
&lt;p&gt;In the same stretch, &lt;a class="link" href="https://aisle.com/blog/curl-adopts-aisle-after-its-ai-agents-discovered-5-cves" target="_blank" rel="noopener"
 &gt;by AISLE&amp;rsquo;s own account&lt;/a&gt;,
an AI security platform contributed 24 pull requests to curl, five of which
earned CVEs, and the project now runs it internally for continuous review. The
same tooling reportedly found &lt;a class="link" href="https://www.lesswrong.com/posts/7aJwgbMEiKq5egQbd/" target="_blank" rel="noopener"
 &gt;all twelve zero-days&lt;/a&gt;
in an OpenSSL release in late January. (Both of those are the tool-makers&amp;rsquo; and a
third party&amp;rsquo;s numbers rather than curl&amp;rsquo;s audited figures, so weigh them as such.
But curl adopting the thing isn&amp;rsquo;t a claim. It&amp;rsquo;s a decision.)&lt;/p&gt;
&lt;p&gt;Sit with the shape of that. curl shut down strangers being paid for AI-shaped
noise, and in the same breath put AI to work as a tool its own maintainers
drive. The two moves look contradictory only if you think &amp;ldquo;AI&amp;rdquo; is a single thing
with a single verdict attached. It isn&amp;rsquo;t. Pointed at the problem by people
accountable for the result, with no prize to farm, it found real bugs. Dangled
in front of anonymous strangers chasing a payout, it produced sand.&lt;/p&gt;
&lt;h2 id="the-tell-is-which-ai-curl-kept-and-which-it-mocked"&gt;The tell is which AI curl kept, and which it mocked
&lt;/h2&gt;&lt;p&gt;Stenberg drew that line about as sharply as a person can. When Anthropic put its
security model, Mythos, in front of curl this spring, it
&lt;a class="link" href="https://daniel.haxx.se/blog/2026/05/11/mythos-finds-a-curl-vulnerability/" target="_blank" rel="noopener"
 &gt;scanned 176,000 lines of C and surfaced a single flaw&lt;/a&gt;,
and Stenberg called the surrounding fanfare
&lt;a class="link" href="https://www.theregister.com/security/2026/05/11/anthropics-bug-hunting-mythos-was-greatest-marketing-stunt-ever-says-curl-creator/5238111" target="_blank" rel="noopener"
 &gt;the greatest marketing stunt he&amp;rsquo;d seen&lt;/a&gt;.
Same maintainer. Adopts one AI, rubbishes another.&lt;/p&gt;
&lt;p&gt;The deciding factor was never whether the thing was AI. Both were. It was
whether the output survived a human checking it, and whether you could check it
at all. AISLE handed over pull requests and CVEs you could read and merge.
Mythos arrived as a closed model and a press release, which is to say a claim
the community has no way to independently test.&lt;/p&gt;
&lt;p&gt;My bias, up front, because it runs the opposite way to what you&amp;rsquo;d expect from
someone writing this: I&amp;rsquo;m a paying Claude subscriber and I lean on Anthropic&amp;rsquo;s
models every working day, the one behind the spadework for this post included.
I&amp;rsquo;m an advocate, not a sceptic, and AI genuinely has its place. That is
&lt;em&gt;exactly&lt;/em&gt; why the Mythos fanfare grates. Overselling a closed model to get out
ahead of the competition, when the one test the public got to see turned up a
single bug, is the sort of thing that chips away at trust in all of it. A result
you can&amp;rsquo;t verify is marketing until proven otherwise, whoever&amp;rsquo;s logo is on the
slide, and I&amp;rsquo;d rather the tools I depend on didn&amp;rsquo;t stoop to it.&lt;/p&gt;
&lt;h2 id="the-cheap-half-and-the-expensive-half"&gt;The cheap half and the expensive half
&lt;/h2&gt;&lt;p&gt;Pull back from curl for a moment, because the lesson isn&amp;rsquo;t really about bounties
at all. Anyone who works with these tools every day knows the same thing: when
they go wrong, it&amp;rsquo;s rarely the model running off on its own. It&amp;rsquo;s the context it
wasn&amp;rsquo;t given, the rope it was handed, the output nobody checked closely enough.
The failure sits on the human side of the keyboard, at the one step that&amp;rsquo;s
easiest to skip, which is verification.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s the pattern curl hit at the scale of an ecosystem. AI made one thing
nearly free: producing work that looks right. It did not make the other thing a
penny cheaper: confirming the work &lt;em&gt;is&lt;/em&gt; right. That cost still falls, in full,
on a person. (A scanner, &lt;a class="link" href="https://phpboyscout.uk/the-security-finding-you-must-not-fix/" &gt;I&amp;rsquo;ve argued before&lt;/a&gt;,
is an argument, not an order; the same goes double for a model.) The bounty&amp;rsquo;s
fatal mistake was paying for the cheap half and quietly assuming it had bought
the expensive one. The same trap waits in code review, in hiring, in CVs read by
machines, but that&amp;rsquo;s a bigger argument for another post.&lt;/p&gt;
&lt;h2 id="pouring-sand-into-the-machine"&gt;Pouring sand into the machine
&lt;/h2&gt;&lt;p&gt;curl didn&amp;rsquo;t capitulate to AI, whatever the headlines decided. It stopped paying
for the worthless half and started using the valuable half, and it had the
discernment to tell a useful tool from a press release while it did so.&lt;/p&gt;
&lt;p&gt;The bounty wasn&amp;rsquo;t a casualty of artificial intelligence. It was a structure
that, the instant plausible output became free, could only fill with sand.
Stenberg said he hopes closing it stops &amp;ldquo;more people pouring sand into the
machine.&amp;rdquo; Reading the last year of his inbox, I think he&amp;rsquo;ll get his wish. The
sand was only ever there because somebody left a bucket of money beside the
funnel.&lt;/p&gt;</content:encoded></item><item><title>House rules</title><link>https://phpboyscout.uk/house-rules/</link><pubDate>Sat, 11 Jul 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/house-rules/</guid><category>ai</category><category>go-tool-base</category><category>Soapbox</category><description>Running three different coding agents against one repository, and the house rules that stopped them undoing each other's work.</description><content:encoded>&lt;p&gt;I&amp;rsquo;ve had a Gemini Pro subscription for about eighteen months. In that time it has
researched a leisure battery and a diesel heater for the campervan conversion,
generated a truly stupid number of pictures of my pets in hats, and helped me build
out a homebrew setting for a D&amp;amp;D campaign that my players still haven&amp;rsquo;t finished.
NotebookLM is genuinely brilliant. The Workspace integration means it can rummage
through my own documents without me copying and pasting half of them into a chat box.
Money well spent.&lt;/p&gt;
&lt;p&gt;&lt;img alt="Three pets in chef’s hats wrecking a kitchen, with a proverb on a sign behind them" class="gallery-image portrait-shot" data-flex-basis="175px" data-flex-grow="73" height="512" loading="lazy" sizes="(max-width: 767px) calc(100vw - 30px), (max-width: 1023px) 700px, (max-width: 1279px) 950px, 1232px" src="https://phpboyscout.uk/house-rules/pets-too-many-cooks_hu_4df2b655612902e4.webp" srcset="https://phpboyscout.uk/house-rules/pets-too-many-cooks_hu_4df2b655612902e4.webp 374w" width="374"&gt;

&lt;/p&gt;
&lt;p&gt;Read the sign on the wall of that kitchen. It says &amp;ldquo;too many cooks on broth&amp;rdquo;. Not
&amp;ldquo;spoil the broth&amp;rdquo;. It&amp;rsquo;s very nearly a proverb, it&amp;rsquo;s rendered beautifully, and it is
wrong, and I didn&amp;rsquo;t notice for a good few seconds because it looked exactly like the
thing it was supposed to be. Hold that thought, because this whole post is about it.&lt;/p&gt;
&lt;p&gt;For writing code, though, it never quite cut the mustard.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s not entirely fair, and it&amp;rsquo;s worth being precise, because there&amp;rsquo;s a version of
this post that&amp;rsquo;s just a man being rude about a product he pays for. About a year ago I
was using the Antigravity IDE, back when I still had a proper desktop to sit at, and it
was massively formative. Not because the code it wrote was better than anything else,
but because it was the first tool that taught me to work &lt;em&gt;with&lt;/em&gt; an agent rather than at
a very enthusiastic autocomplete. That shift, from &amp;ldquo;finish my line&amp;rdquo; to &amp;ldquo;here&amp;rsquo;s a task,
go and do it, tell me what you did&amp;rdquo;, is the whole game. I learned that in Antigravity.&lt;/p&gt;
&lt;p&gt;Then I gave the work laptop back. The new job&amp;rsquo;s replacement laptop isn&amp;rsquo;t mine to do
personal work on, quite rightly, so my primary environment became a headless dev server
I reach over SSH. No desktop, no IDE, no window to drag things around in. And in a
terminal, Claude Code turned out to be the thing that fit. It became my daily driver
almost by accident, having barely used it before.&lt;/p&gt;
&lt;p&gt;So: a Gemini subscription I still pay for and use every week, and a Claude subscription
doing the actual engineering. Two tools, cleanly separated, no conflict.&lt;/p&gt;
&lt;p&gt;Then I ran out of tokens.&lt;/p&gt;
&lt;h2 id="reaching-for-what-youve-already-paid-for"&gt;Reaching for what you&amp;rsquo;ve already paid for
&lt;/h2&gt;&lt;p&gt;A week without quota is a long time. Long enough that I stopped waiting for the reset
and looked at what was already sitting in my account, which is when I noticed Google had
shipped agy, a CLI in the same shape as Claude Code, riding the subscription I&amp;rsquo;d been
paying for since before I owned a campervan.&lt;/p&gt;
&lt;p&gt;It was disappointing. It made mistakes. Some were the sort you catch on the next line,
and some sat there waiting until I bothered to look.&lt;/p&gt;
&lt;p&gt;Now some of that is the models. Gemini gives me two that matter, 3.5 Flash and 3.1 Pro,
and Google will happily tell you 3.5 Flash is the strongest coding model they&amp;rsquo;ve got.
Read that again though. The strongest &lt;em&gt;they&amp;rsquo;ve got&lt;/em&gt;. It beats 3.1 Pro, which is a very
different sentence from &amp;ldquo;it beats what everybody else shipped this month&amp;rdquo;. Coding has
never been where Gemini lives.&lt;/p&gt;
&lt;p&gt;And some of it is me. I&amp;rsquo;d picked up a brand new CLI and driven it exactly like the one I
already knew, because it looked like the one I already knew, and at no point did I open
its documentation to find out what it could actually do or how to point it at the right
model for the job I was giving it. RTFM. I have said that to other people, out loud, more
than once. I got round to taking my own advice about a week later, by which time I&amp;rsquo;d spent
a fair while blaming the tool for a decent share of my own idleness.&lt;/p&gt;
&lt;p&gt;Lesson learned. Again.&lt;/p&gt;
&lt;p&gt;Where it does earn its keep, mind, is ideation. I set it loose on
&lt;a class="link" href="https://krites.phpboyscout.uk" target="_blank" rel="noopener"
 &gt;krites&lt;/a&gt;, my photo culling tool, and told it to dream up
features and redesign the interface. I wasn&amp;rsquo;t after accuracy.
I wanted creative flair, and Gemini&amp;rsquo;s multimodal chops made that quick and clean. It came
back with thirty-three feature ideas and eight interface concepts in an evening. Most of
them were wrong. Several very much weren&amp;rsquo;t. That&amp;rsquo;s a respectable evening by anyone&amp;rsquo;s
measure, and it&amp;rsquo;s the job agy does around here now.&lt;/p&gt;
&lt;p&gt;Which left me with no tokens for Claude, and an agent I&amp;rsquo;d just adopted that couldn&amp;rsquo;t move
the actual projects along. So I had a punt on codex. Twenty dollars, on the strength of
what people had been saying about GPT-5.5, and because Anthropic&amp;rsquo;s Max plan is a number I
can&amp;rsquo;t look at with a straight face for what is, after all, a hobby. It was the right call.
Claude is still the daily driver, tokens permitting. But codex is a proper fallback now,
and more useful than that, a second opinion. It reviews Claude&amp;rsquo;s work. Claude reviews its
work right back.&lt;/p&gt;
&lt;p&gt;Three agents. One repository.&lt;/p&gt;
&lt;h2 id="three-guests-one-hallway"&gt;Three guests, one hallway
&lt;/h2&gt;&lt;p&gt;The problem, when it came, had nothing to do with the models.&lt;/p&gt;
&lt;p&gt;Every one of these tools wants to be told how the project works. What the commit
conventions are. That we use &lt;code&gt;just&lt;/code&gt; and not &lt;code&gt;make&lt;/code&gt;. That specs live in
&lt;code&gt;docs/development/specs/&lt;/code&gt; and you don&amp;rsquo;t write code until one is approved. That there is
no AI attribution in commit messages, ever, because the human who approves the commit
owns it entirely. That we&amp;rsquo;re pre-1.0, so a &lt;code&gt;BREAKING CHANGE:&lt;/code&gt; footer will cause the
release automation to cheerfully cut a v1 we are not ready for.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s 243 lines of hard-won house style. I had it in a &lt;code&gt;CLAUDE.md&lt;/code&gt;, where Claude reads
it. And now two more agents had walked through the door with no idea about any of it.&lt;/p&gt;
&lt;p&gt;You know the sign. The one in the front hallway of a certain kind of house, usually
wooden, usually in a font somebody&amp;rsquo;s aunt chose.&lt;/p&gt;

 &lt;blockquote&gt;
 &lt;ol&gt;
&lt;li&gt;Take off your shoes&lt;/li&gt;
&lt;li&gt;Clean up after yourself&lt;/li&gt;
&lt;li&gt;Mum is always right&lt;/li&gt;
&lt;/ol&gt;

 &lt;/blockquote&gt;
&lt;p&gt;Nobody hands each guest a personalised laminated rulebook on the doorstep. There&amp;rsquo;s one
sign, everyone reads it, and the rules are the rules whether you&amp;rsquo;re family or you&amp;rsquo;ve
come to fix the boiler. Standardisation is the only thing that makes a house with guests
in it survivable.&lt;/p&gt;
&lt;p&gt;So the instructions came out of &lt;code&gt;CLAUDE.md&lt;/code&gt; and went into an &lt;code&gt;AGENTS.md&lt;/code&gt;, which is the
name codex and agy both look for without being asked. The sign went up in the hallway.&lt;/p&gt;
&lt;h2 id="the-adapter-is-not-a-wart"&gt;The adapter is not a wart
&lt;/h2&gt;&lt;p&gt;My first instinct was that &lt;code&gt;CLAUDE.md&lt;/code&gt; should now die. One file, one truth, and Anthropic
should get with the programme and read &lt;code&gt;AGENTS.md&lt;/code&gt; like everybody else.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve changed my mind, and the reason is sitting in the same repository, in a completely
different refactor, that I was working on the same week.&lt;/p&gt;
&lt;p&gt;We&amp;rsquo;re pulling &lt;code&gt;go-tool-base&lt;/code&gt; apart. Not breaking it up, just loosening it, so the genuinely
reusable bits can leave without dragging the entire framework out of the door behind them.
So each package now owns a little typed struct that says, plainly, here is what I need in
order to run: a timeout, an endpoint, a token. And the framework keeps a small file sitting
next to it, &lt;code&gt;config_adapter.go&lt;/code&gt;, whose whole job is to take the framework&amp;rsquo;s own sprawling
config object and fill that struct in. The package holds the shared truth. The adapter is
thin, it&amp;rsquo;s local, and it exists &lt;em&gt;precisely so the package never has to know the framework
is there at all&lt;/em&gt;.&lt;/p&gt;
&lt;p&gt;Look again at what &lt;code&gt;CLAUDE.md&lt;/code&gt; is now. &lt;code&gt;AGENTS.md&lt;/code&gt; holds the shared truth. &lt;code&gt;CLAUDE.md&lt;/code&gt; is
a small local file that exists so I can give Claude an instruction the other two would
only find confusing, without that instruction leaking into the shared file. Codex has
somewhere to put its quirks. agy has somewhere to put its quirks.&lt;/p&gt;
&lt;p&gt;It&amp;rsquo;s the same shape. Extract the core, leave a thin adapter at the border.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;d love to tell you I saw that symmetry coming and designed it in. I didn&amp;rsquo;t. My head was
in that space already, and the dots joined themselves while I was thinking about
something else. That&amp;rsquo;s usually how it goes, and I&amp;rsquo;d rather admit it than pretend to a
grand plan.&lt;/p&gt;
&lt;p&gt;So no, the per-agent file isn&amp;rsquo;t a wart. It&amp;rsquo;s the adapter, and the other vendors should
consider growing one.&lt;/p&gt;
&lt;h2 id="the-sign-nobody-read"&gt;The sign nobody read
&lt;/h2&gt;&lt;p&gt;There is a hole in this, and it&amp;rsquo;s mine.&lt;/p&gt;
&lt;p&gt;I asked agy to leave &lt;code&gt;CLAUDE.md&lt;/code&gt; behind as a pointer to &lt;code&gt;AGENTS.md&lt;/code&gt;, with a file
reference include, which I was fairly sure was possible. What it wrote was this:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-markdown" data-lang="markdown"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;&amp;gt; &lt;/span&gt;&lt;span class="ge"&gt;**Note:** The core agent instructions have been consolidated for use across all our
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;&amp;gt; &lt;/span&gt;&lt;span class="ge"&gt;AI tools (Claude, agy, codex).
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;&amp;gt; &lt;/span&gt;&lt;span class="ge"&gt;Please read and follow the instructions in [`AGENTS.md`](AGENTS.md) for all general
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="k"&gt;&amp;gt; &lt;/span&gt;&lt;span class="ge"&gt;development workflows, architecture guidelines, and commit conventions.
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Which looks fine. Reads fine. It is, in fact, a polite request in prose with a hyperlink
attached, and it is not an include of anything.&lt;/p&gt;
&lt;p&gt;Too many cooks on broth.&lt;/p&gt;
&lt;p&gt;Claude Code reads &lt;code&gt;CLAUDE.md&lt;/code&gt;. It does not read &lt;code&gt;AGENTS.md&lt;/code&gt;, and a markdown link is not a
loading instruction, it&amp;rsquo;s a hint the model may or may not act on. There &lt;em&gt;is&lt;/em&gt; a real import
syntax, a bare &lt;code&gt;@AGENTS.md&lt;/code&gt; on its own line, which expands the file into context when the
session starts. I&amp;rsquo;d guessed right that it existed. It was documented the whole time. The
agent didn&amp;rsquo;t use it, and I didn&amp;rsquo;t check.&lt;/p&gt;
&lt;p&gt;RTFM, it turns out, is a lesson you get to learn twice in the same fortnight.&lt;/p&gt;
&lt;p&gt;So for three days, agy and codex walked into the hallway and read all 243 lines of the
sign, and Claude, the agent the original file was named after, walked in and read eight
lines of a note telling it there was a sign somewhere.&lt;/p&gt;
&lt;p&gt;I proved it in the end by asking a fresh session for a string that only exists in
&lt;code&gt;AGENTS.md&lt;/code&gt;. With the link: not found. With the import: found. One line, one merge
request, and the guest can see the sign.&lt;/p&gt;
&lt;p&gt;Now, I&amp;rsquo;ve written before that
&lt;a class="link" href="https://phpboyscout.uk/the-interpreter-we-forgot-to-sandbox/" &gt;a &lt;code&gt;CLAUDE.md&lt;/code&gt; is source code and the agent is its interpreter&lt;/a&gt;.
I meant it as a warning about what an attacker could put in one. It cuts the other way
too. If it&amp;rsquo;s source code, then a link where an import belongs isn&amp;rsquo;t a typo in a document,
it&amp;rsquo;s a missing &lt;code&gt;import&lt;/code&gt; statement at the top of a file. The program still runs. It just
doesn&amp;rsquo;t have the thing it needed, and nothing tells you.&lt;/p&gt;
&lt;h2 id="only-human"&gt;Only human
&lt;/h2&gt;&lt;p&gt;Whose fault?&lt;/p&gt;
&lt;p&gt;Mine, of course. Not agy&amp;rsquo;s for writing plausible markdown. Not Anthropic&amp;rsquo;s, though I&amp;rsquo;d
happily take an &lt;code&gt;AGENTS.md&lt;/code&gt; read by default.&lt;/p&gt;
&lt;p&gt;I have been caught out before by an agent&amp;rsquo;s confidence in its own ability, and I&amp;rsquo;ll be
caught out again. It&amp;rsquo;s the same way you get caught out by a keen junior engineer: they
tell you it&amp;rsquo;s done, they believe it&amp;rsquo;s done, they have every reason to think it&amp;rsquo;s done, and
they are wrong in a way that only shows up later. The failure isn&amp;rsquo;t theirs. The failure is
that I didn&amp;rsquo;t go and check the maths.&lt;/p&gt;
&lt;p&gt;Except a junior is
&lt;a class="link" href="https://phpboyscout.uk/the-rung-we-sawed-off/" &gt;a senior who hasn&amp;rsquo;t happened yet&lt;/a&gt;.
Give one two years and they&amp;rsquo;ll be checking &lt;em&gt;your&lt;/em&gt; maths. This one won&amp;rsquo;t. Next week it
will make the same class of mistake with exactly as much confidence, and the only thing
standing between that and my &lt;code&gt;main&lt;/code&gt; branch is whether I could be bothered to look.&lt;/p&gt;
&lt;p&gt;Three agents in tandem takes a lot of oversight. There&amp;rsquo;s a lot of knowledge that has to
move between them, and unguarded mistakes can and will happen. Standardisation is the
only lever I&amp;rsquo;ve got, at least until the providers agree a common standard between
themselves. Until then I make do, if only to avoid typing out the same instructions for
every guest who walks in.&lt;/p&gt;
&lt;h2 id="the-thirteen-i-couldnt-import"&gt;The thirteen I couldn&amp;rsquo;t import
&lt;/h2&gt;&lt;p&gt;I&amp;rsquo;ll leave you with the bit I haven&amp;rsquo;t solved.&lt;/p&gt;
&lt;p&gt;The sign in the hallway says &amp;ldquo;read the house rules&amp;rdquo;, but it also says &amp;ldquo;and the detailed
procedures are through there&amp;rdquo;. Those procedures are skills: how to draft a spec, how to
verify before a merge request, how to write a conventional commit. I already publish them,
as a plugin marketplace, because I&amp;rsquo;d rather write them once.&lt;/p&gt;
&lt;p&gt;Claude can install them from there. agy and codex can&amp;rsquo;t reach it. So there are now
fourteen skill files sitting inside the repository, thirteen of which are &lt;em&gt;copies&lt;/em&gt; of
something I already publish somewhere else, checked in so that the other two guests can
read them.&lt;/p&gt;
&lt;p&gt;I&amp;rsquo;ve spent a fortnight taking a framework apart so that its packages could be imported
instead of carried around. And then, one floor up, I copied thirteen files into a repo
because I had no way to import them.&lt;/p&gt;
&lt;p&gt;I haven&amp;rsquo;t the faintest idea what the fix looks like yet.&lt;/p&gt;</content:encoded></item><item><title>Technical CV writing is still hard, and now a robot reads it first</title><link>https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/</link><pubDate>Fri, 22 May 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/</guid><category>career-meta</category><category>Campfire</category><description>My CV stopped landing because a machine reads it before a person does. Reworking the structure for the filter without losing the voice.</description><content:encoded>&lt;p&gt;Seven years ago I wrote a post called &lt;a class="link" href="https://phpboyscout.uk/technical-cv-writing/" &gt;Technical CV writing is hard&lt;/a&gt;, pulled my own CV apart, and explained every choice in it. I even bragged that it converted to a first interview about eighty per cent of the time, then added &amp;ldquo;watch me now jinx myself for the future&amp;rdquo;. Reader, I jinxed myself. I&amp;rsquo;m back on the market, the same CV that served me for two decades went out into the world, and what came back was a sort of stunned silence. Not even rejections. Just nothing.&lt;/p&gt;
&lt;h2 id="the-cv-that-suddenly-stopped-working"&gt;The CV that suddenly stopped working
&lt;/h2&gt;&lt;p&gt;The thing about that silence is how &lt;em&gt;specific&lt;/em&gt; it was. Some applications behaved exactly as they always had: a human read the CV, liked it or didn&amp;rsquo;t, and replied like a person. Others went into a void. And the void had a pattern to it. It was the bigger, more process-heavy outfits, the ones you&amp;rsquo;d bet good money have an Applicant Tracking System and an &amp;ldquo;AI-assisted screening&amp;rdquo; line item in some HR budget.&lt;/p&gt;
&lt;p&gt;That&amp;rsquo;s when the penny dropped. My CV wasn&amp;rsquo;t failing to impress anyone. It wasn&amp;rsquo;t reaching anyone. The first thing reading it wasn&amp;rsquo;t a person at all.&lt;/p&gt;
&lt;h2 id="the-reader-changed-and-i-hadnt-noticed"&gt;The reader changed, and I hadn&amp;rsquo;t noticed
&lt;/h2&gt;&lt;p&gt;I&amp;rsquo;ve made this exact point on this blog before, only about software: &lt;a class="link" href="https://phpboyscout.uk/half-your-users-dont-have-eyes/" &gt;half your users don&amp;rsquo;t have eyes&lt;/a&gt;. A CLI tool&amp;rsquo;s output has two audiences, the human at the terminal and the script parsing the output, and they want completely different things. It turns out a CV is now precisely the same. It has two readers, and the first one is a machine.&lt;/p&gt;
&lt;p&gt;A human recruiter reads a CV the way I designed mine to be read: narrative, personality, a sense of the person. An ATS or an AI screen does nothing of the sort. It parses for structure, for keyword density, for recency, for numbers it can latch onto. My CV was a beautifully tailored sales pitch aimed squarely at a human who, increasingly, never gets to see it, because a parser in front of them scored it and quietly binned it first.&lt;/p&gt;
&lt;p&gt;Everything that made it a good &lt;em&gt;human&lt;/em&gt; document was, to the machine, either invisible or actively confusing.&lt;/p&gt;
&lt;h2 id="so-i-asked-an-ai-what-the-ai-hated"&gt;So I asked an AI what the AI hated
&lt;/h2&gt;&lt;p&gt;There&amp;rsquo;s an irony here I&amp;rsquo;m choosing to enjoy rather than resent. The way I worked out what the filters object to was to sit down with Gemini, hand it my CV, and ask it to read the thing the way a recruitment AI would and tell me where it tripped. Using one AI to get past another. Fight fire with fire.&lt;/p&gt;
&lt;p&gt;The one instruction I was firm about, and I&amp;rsquo;ll come back to it, was that the CV had to stay recognisably &lt;em&gt;me&lt;/em&gt;. I wasn&amp;rsquo;t asking Gemini to launder my career into something generic and machine-shaped. I was asking it to help me keep as much of my own voice and judgement as possible, while making the thing easier for an AI to approve and a human to enjoy. There&amp;rsquo;s a practical edge to that, too: the screening tools are increasingly tuned to spot the patterns of generated text and weight them down, so a CV that reads as though a model wrote it can trip the very filter you were trying to please, quite apart from leaving the human at the end of it cold.&lt;/p&gt;
&lt;p&gt;With that ground rule set, the hurdles it surfaced were genuinely illuminating, and a bit humbling given I&amp;rsquo;d written a whole confident blog post about how to do this.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;The skills tables are worse than useless.&lt;/strong&gt; My CV led with two lovely tables: Management Skills and Technical Skills, each with a level and years of experience. Clean and scannable for a human. To a lot of parsers, a table is a trap: they flatten it into a jumble and lose the structure entirely. Worse, listing &amp;ldquo;20+ years&amp;rdquo; against nearly everything triggers what I can only call the recency trap. Modern screening looks for skills that show up &lt;em&gt;in your recent job descriptions&lt;/em&gt;, not in a header table. A language sitting in my skills table but not in my last two roles reads as stale or unverified, no matter how many years I claimed next to it. Gemini put it plainly: &amp;ldquo;if a tool sees Golang in a top table but doesn&amp;rsquo;t see it explicitly mentioned in your last two job descriptions, it assumes the skill is stale or unverified.&amp;rdquo;&lt;/p&gt;
&lt;p&gt;&lt;img alt="My skills laid out as tables of skill, level and commercial experience. Lovely for a human to scan, a jumble the moment a parser flattens the formatting. This is the long-standing shape, here in its original 2019 form." class="gallery-image" data-flex-basis="251px" data-flex-grow="104" height="792" loading="lazy" sizes="(max-width: 767px) calc(100vw - 30px), (max-width: 1023px) 700px, (max-width: 1279px) 950px, 1232px" src="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-before_hu_5ee9e1318c5d7ff5.webp" srcset="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-before_hu_50bbba9ea10ac31a.webp 480w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-before_hu_277ffc3e5e12239d.webp 720w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-before_hu_5ee9e1318c5d7ff5.webp 831w" width="831"&gt;

&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;&amp;ldquo;I have a passion for what I do&amp;rdquo; is noise.&lt;/strong&gt; My opening profile statement, which I was rather proud of, is exactly the sort of thing a screening tool discards wholesale. As Gemini noted, these tools &amp;ldquo;completely ignore subjective self-assessments &amp;hellip; because they cannot be measured or verified.&amp;rdquo; It wants a dense, factual summary full of the nouns it&amp;rsquo;s searching for, right at the top.&lt;/p&gt;
&lt;p&gt;&lt;img alt="The old opening: my name, my contact details, and a warm but entirely unmeasurable “I have a passion for what I do” profile statement." class="gallery-image" data-flex-basis="1055px" data-flex-grow="439" height="188" loading="lazy" sizes="(max-width: 767px) calc(100vw - 30px), (max-width: 1023px) 700px, (max-width: 1279px) 950px, 1232px" src="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-before_hu_f5f3517119e92c75.webp" srcset="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-before_hu_5343d7503df312a5.webp 480w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-before_hu_6164b92783de19e5.webp 720w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-before_hu_f5f3517119e92c75.webp 827w" width="827"&gt;

&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;My numbers thin out the further back you go.&lt;/strong&gt; My recent roles are full of the data these tools love: a 75% reduction in deployment times, three thousand-odd Kubernetes clusters, a GitLab instance with four hundred thousand repositories. My older roles, written years ago in a more narrative style, are all &amp;ldquo;oversaw the delivery of solutions&amp;rdquo; with not a metric in sight. The machine reads that as a career that got vaguer over time, which is the opposite of true.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Four pages is at least two too many.&lt;/strong&gt; Parsers weight the first page or two most heavily. My education and the foundational stuff sat on pages three and four, where the algorithm barely bothers to look.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;It couldn&amp;rsquo;t work out what I am.&lt;/strong&gt; This was the sharp one. With &amp;ldquo;pre-sales&amp;rdquo;, &amp;ldquo;client management&amp;rdquo; and &amp;ldquo;Managing Director&amp;rdquo; sitting next to deep technical keywords, the classifier genuinely can&amp;rsquo;t decide whether I&amp;rsquo;m a commercial manager who used to code or a hands-on engineer who drifted into management. As Gemini described it: &amp;ldquo;the algorithm gets confused &amp;hellip; It struggles to classify you: Are you a commercial manager who used to code, or a hands-on techie who got pushed into management?&amp;rdquo; So it does the safe thing and matches me to neither.&lt;/p&gt;
&lt;h2 id="what-im-actually-changing"&gt;What I&amp;rsquo;m actually changing
&lt;/h2&gt;&lt;p&gt;Knowing the hurdles, here&amp;rsquo;s what the rebuild looks like. This is the part I want to be useful, so it&amp;rsquo;s concrete.&lt;/p&gt;
&lt;p&gt;The tables are gone. In their place is a &amp;ldquo;Core Expertise&amp;rdquo; section, plain text the parser can read, grouped so my leadership sits next to my technical stack. And I&amp;rsquo;ve done the thing 2019-me was too much of a show-off to do: tiered it &lt;em&gt;honestly&lt;/em&gt;. Instead of &amp;ldquo;Expert+&amp;rdquo; against everything, there&amp;rsquo;s a primary tier of what I actually do day to day, a proficient tier I can deploy without blinking, and a frank &amp;ldquo;familiar, not current&amp;rdquo; tier for the languages I last touched in anger a decade ago. That honesty isn&amp;rsquo;t just decency. A wall of &amp;ldquo;expert at everything&amp;rdquo; reads as noise to a machine and as bluster to a human, and I&amp;rsquo;d been doing both.&lt;/p&gt;
&lt;p&gt;&lt;img alt="The replacement: a plain-text Core Expertise list a parser can read straight through, tiered honestly into what I do day to day, what I’m proficient in, and what I’m only still familiar with." class="gallery-image" data-flex-basis="288px" data-flex-grow="120" height="1149" loading="lazy" sizes="(max-width: 767px) calc(100vw - 30px), (max-width: 1023px) 700px, (max-width: 1279px) 950px, 1232px" src="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-after_hu_61e84b5e406d71d4.webp" srcset="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-after_hu_3377bae072cc3175.webp 480w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-after_hu_abfe24012095612f.webp 720w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-after_hu_1df16e2dd124629d.webp 1080w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-skills-after_hu_61e84b5e406d71d4.webp 1383w" width="1383"&gt;

&lt;/p&gt;
&lt;p&gt;The subjective profile is replaced with a keyword-rich professional summary that says, in the first two lines, exactly what I am and at what scale.&lt;/p&gt;
&lt;p&gt;&lt;img alt="The replacement: a Professional Summary that leads with the role and the scale, in the nouns a parser is actually hunting for, with the person still audible underneath." class="gallery-image" data-flex-basis="623px" data-flex-grow="259" height="537" loading="lazy" sizes="(max-width: 767px) calc(100vw - 30px), (max-width: 1023px) 700px, (max-width: 1279px) 950px, 1232px" src="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-after_hu_2312f05c7f8b99bb.webp" srcset="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-after_hu_a068217fa87974ea.webp 480w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-after_hu_6f974cc746d43525.webp 720w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-after_hu_7d8e4402f878913d.webp 1080w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-profile-after_hu_2312f05c7f8b99bb.webp 1394w" width="1394"&gt;

 The keywords that mattered have been woven down &lt;em&gt;into&lt;/em&gt; the recent role bullets, so the parser sees them where it trusts them. And I&amp;rsquo;ve reframed the people-management and pre-sales language toward technical enablement and architectural advisory, because what I&amp;rsquo;m actually chasing is the technical-leader sweet spot: the person who owns the architecture and mentors the engineers, without the HR admin and the sales pitches. The CV now points at that, deliberately, so the classifier stops dithering.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s also a more personal beat in here. A previous employer handed me a role with a &amp;ldquo;VP&amp;rdquo; title, sold to me as exactly the technical-leadership job I&amp;rsquo;d been chasing. It wasn&amp;rsquo;t. The title turned out to be a pay-grade bracket rather than a description of the work, the work itself was hands-on firefighting with little of the leadership or empowerment I&amp;rsquo;d been promised, and I moved on within a few months. To a screening AI, that pairing is doubly awkward. A &amp;ldquo;VP&amp;rdquo; title files me as a meeting-heavy executive and rules me out of the hands-on Principal and Lead roles I actually want, and a sub-six-month stint trips the flight-risk flag that some trackers quietly score you down for. So the fix is to stop letting the inflated label do the talking: describe the functional reality of the work, retitle it to the technical track it actually was, and let the scale of what I wrestled with speak instead of the job title. Titles, it turns out, are for the pay band. The bullets are for the truth.&lt;/p&gt;
&lt;p&gt;&lt;img alt="My recent roles on the new CV: each leads with the work and the numbers, in technical-track titles a parser weights and a human believes." class="gallery-image" data-flex-basis="282px" data-flex-grow="117" height="1167" loading="lazy" sizes="(max-width: 767px) calc(100vw - 30px), (max-width: 1023px) 700px, (max-width: 1279px) 950px, 1232px" src="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-experience_hu_e68702fb97f43fbe.webp" srcset="https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-experience_hu_74269b608a2268a6.webp 480w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-experience_hu_7757bb28ec3010e3.webp 720w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-experience_hu_8865bec496bcd53d.webp 1080w, https://phpboyscout.uk/technical-cv-writing-and-the-ai-filter/cv-experience_hu_e68702fb97f43fbe.webp 1374w" width="1374"&gt;

&lt;/p&gt;
&lt;h2 id="keeping-myself-in-it"&gt;Keeping myself in it
&lt;/h2&gt;&lt;p&gt;Back to that ground rule. Every one of these changes is in service of getting past the machine to the human behind it, and neither reader is well served by a CV with the person scrubbed out of it. The screen, increasingly, is trained to notice generic generated phrasing and mark it down; the human, always, would rather read something with a pulse. So the keywords go in, the structure gets fixed, the metrics come forward, and the &lt;em&gt;voice stays mine&lt;/em&gt;. No &amp;ldquo;results-driven synergistic leveraging of cross-functional paradigms&amp;rdquo; that nobody would ever say out loud. That was the whole point of doing it this way: let the AI help reshape the &lt;em&gt;structure&lt;/em&gt; a parser cares about, while the &lt;em&gt;words&lt;/em&gt; stay mine, so what comes out is easier for a machine to approve, easier for a human to enjoy, and still unmistakably written by me. Optimising for the filter and sounding like myself turned out not to be in conflict at all.&lt;/p&gt;
&lt;h2 id="i-genuinely-dont-know-if-this-works-yet"&gt;I genuinely don&amp;rsquo;t know if this works yet
&lt;/h2&gt;&lt;p&gt;Here&amp;rsquo;s the part that makes this a post and not a victory lap. I don&amp;rsquo;t know if any of this lands. The old CV converted at around eighty per cent, on my own possibly-generous reckoning, right up until it abruptly didn&amp;rsquo;t. The new one is going out now, into the same market and the same filters that were stonewalling me a fortnight ago.&lt;/p&gt;
&lt;p&gt;So this is a promise as much as a post. I&amp;rsquo;m going to keep count, the way I should have all along, and come back with the actual numbers: did reshaping my CV for a reader with no eyes genuinely move the needle, or did I just make it uglier and learn nothing? Either way you&amp;rsquo;ll get the truth, because a follow-up that only reports good news isn&amp;rsquo;t worth writing. Watch this space, and if you&amp;rsquo;re sending CVs into the same silence, maybe try reading yours the way a machine would first. It&amp;rsquo;s a deeply odd exercise, and I suspect it&amp;rsquo;s now an essential one.&lt;/p&gt;</content:encoded></item><item><title>About that 'no AI'...</title><link>https://phpboyscout.uk/about-that-no-ai/</link><pubDate>Wed, 15 Jul 2026 00:00:00 +0000</pubDate><guid isPermaLink="true">https://phpboyscout.uk/about-that-no-ai/</guid><category>krites</category><category>Pioneering</category><description>A correction to my own post: the culler does contain machine learning, and the distinction I drew between ML and AI does not survive scrutiny.</description><content:encoded>&lt;p&gt;A couple of weeks back I made a claim I was happy to put my name to: &lt;a class="link" href="https://phpboyscout.uk/no-ai-in-my-photo-culler/" &gt;There&amp;rsquo;s no AI in my photo culler&lt;/a&gt;. And I meant it. The first pass that sorts Hailey&amp;rsquo;s wedding photos into keep, maybe and reject is arithmetic, the sort a patient enough person could do by hand. Blur is the variance of a Laplacian, a duplicate is sixty-four bits that happen to agree, exposure is a histogram. Not a trained weight in sight.&lt;/p&gt;
&lt;p&gt;Then I went and put two neural networks in it.&lt;/p&gt;
&lt;p&gt;So&amp;hellip; about that. When I said &amp;ldquo;no AI&amp;rdquo;, what I really meant was &amp;ldquo;no &lt;em&gt;generative&lt;/em&gt; AI&amp;rdquo;, and the gap between those two has turned out to be the whole point. There&amp;rsquo;s machine learning in krites now, there&amp;rsquo;s more on the way, and &amp;ldquo;AI&amp;rdquo; has become a word that can&amp;rsquo;t tell the two apart.&lt;/p&gt;
&lt;h2 id="the-bit-i-always-said-was-coming"&gt;The bit I always said was coming
&lt;/h2&gt;&lt;p&gt;In that culler post I drew a line. The heavy lifting of the first pass would always be model-free, I said, but some of what a photographer culls on genuinely needs a model: is this person mid-blink, is anyone actually looking at the camera, is the composition any good. Those would be model-backed when they came, and they&amp;rsquo;d sit outside the deterministic core, behind an interface, opt-in. That was a promissory note. Well, the first of them has come due: the eye-and-blink signal has landed, and it pays out exactly as I said it would.&lt;/p&gt;
&lt;p&gt;Spotting a blink genuinely needs a model. To know whether someone&amp;rsquo;s eyes are open you first have to find the face, then find the eyes on it, and there&amp;rsquo;s no way to do that with a &lt;code&gt;for&lt;/code&gt; loop over pixel brightness. So krites does it with two models, both run through ONNX Runtime: UltraFace to find the faces, and InsightFace&amp;rsquo;s 106-point landmark model to map the features on each one (&lt;a class="link" href="https://gitlab.com/phpboyscout/krites/-/blob/67a66ad/pkg/face/models/models.go#L31-L44" target="_blank" rel="noopener"
 &gt;the pinned models&lt;/a&gt;). Both are neural networks, trained on faces. This is machine learning, the thing the first pass pointedly did without. It&amp;rsquo;s off by default, it lives in its own package, and it sits behind the same &lt;code&gt;face.Analyzer&lt;/code&gt; interface the post promised, so the deterministic engine never imports it (spec &lt;a class="link" href="https://gitlab.com/phpboyscout/krites/-/blob/67a66ad/docs/development/specs/0004-face-eye.md" target="_blank" rel="noopener"
 &gt;&lt;code&gt;0004&lt;/code&gt;&lt;/a&gt;).&lt;/p&gt;
&lt;h2 id="but-look-at-what-it-gives-back"&gt;But look at what it gives back
&lt;/h2&gt;&lt;p&gt;Here&amp;rsquo;s why two neural networks still don&amp;rsquo;t add up to the kind of &amp;ldquo;AI&amp;rdquo; everyone pictures now. The models don&amp;rsquo;t &lt;em&gt;decide&lt;/em&gt; anything. They hand back coordinates: a box around each face, and a hundred-odd points marking its features. The blink call is then plain geometry on six of those points per eye, the eye-aspect-ratio, the distance across the eyelids over the distance across the eye, a number that falls towards zero as the lid closes:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-go" data-lang="go"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// eyeAspectRatio is the Soukupová–Čech eye-aspect-ratio for a 6-point eye&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// contour (p0..p5 ordered around the eye, starting at the outer corner):&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;//
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;//	EAR = (|p1-p5| + |p2-p4|) / (2 · |p0-p3|)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;//
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// the two vertical lid distances over twice the horizontal width. Higher means&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// more open; a blink collapses the verticals toward zero.&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;I calibrated the open and closed thresholds against 556 frames from &lt;a class="link" href="https://phpboyscout.uk/i-built-my-wife-a-judge/" &gt;the wedding that started this whole project&lt;/a&gt;, actual blinks and actual open eyes, so the numbers mean something for the job (&lt;a class="link" href="https://gitlab.com/phpboyscout/krites/-/blob/67a66ad/pkg/face/onnx/config.go#L46-L60" target="_blank" rel="noopener"
 &gt;the anchors&lt;/a&gt;). And the interface it plugs into doesn&amp;rsquo;t just hope it behaves itself, it demands it:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre tabindex="0" class="chroma"&gt;&lt;code class="language-go" data-lang="go"&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// Analyzer finds faces in a frame and reports per-face eye state. Implementations&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// are model-backed (pkg/face/onnx) or faked in tests; the engine codes only to&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// this interface (0004 R-FACE-3). Analyze MUST be deterministic for a given&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// (img, model) and honour ctx cancellation/timeout (0004 R-FACE-4; 0002&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="c1"&gt;// R-GLOBAL-7, R-GLOBAL-8).&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="kd"&gt;type&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;Analyzer&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kd"&gt;interface&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;{&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="w"&gt;	&lt;/span&gt;&lt;span class="nf"&gt;Analyze&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;ctx&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;context&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Context&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;img&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;image&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="nx"&gt;Image&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nx"&gt;Result&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kt"&gt;error&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;span class="line"&gt;&lt;span class="cl"&gt;&lt;span class="p"&gt;}&lt;/span&gt;&lt;span class="w"&gt;
&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;&lt;p&gt;Same frame in, same answer out, every time (&lt;a class="link" href="https://gitlab.com/phpboyscout/krites/-/blob/67a66ad/pkg/analyze/face/face.go#L19-L27" target="_blank" rel="noopener"
 &gt;&lt;code&gt;face.go&lt;/code&gt;&lt;/a&gt;). That one requirement is the line between this and a chatbot. It&amp;rsquo;s machine learning that behaves like the arithmetic culler it sits next to: reproducible, debuggable, and tunable by Hailey rather than by me or a vendor. The model finds the eyes; a number she can move decides what counts as shut. And a blink only ever demotes a frame to &lt;em&gt;maybe&lt;/em&gt;, it never rejects one on its own, because krites proposes and the human disposes.&lt;/p&gt;
&lt;h2 id="the-word-got-narrow"&gt;The word got narrow
&lt;/h2&gt;&lt;p&gt;Now the actual argument. Ask most people what &amp;ldquo;AI&amp;rdquo; is today and they&amp;rsquo;ll describe a chat box. You type, it types back; it writes, paints, codes. Generative, conversational, and genuinely astonishing. I use it every day. But it&amp;rsquo;s one slice of a very large pie, and somewhere in the last three or four years it ate the whole word.&lt;/p&gt;
&lt;p&gt;There&amp;rsquo;s no &amp;ldquo;versus&amp;rdquo; in any of this. Generative models and the quiet predictive kind are branches of the same family, and I&amp;rsquo;m a fan of the lot. The point is only that &amp;ldquo;AI&amp;rdquo; used to name the whole field and now names one corner of it. Machine learning is enormous, decades older than the chatbot, and the overwhelming majority of it has nothing to say to you at all. It takes a varied input and returns a consistent, predictable output. It finds a face. It reads a postcode off an envelope. It flags the dodgy transaction. It spots the tumour on the scan. None of it talks, none of it improvises, and almost none of it gets called &amp;ldquo;AI&amp;rdquo; any more.&lt;/p&gt;
&lt;p&gt;Some of the most formative work of my career, years before GPT was a name anyone outside the field would recognise, was built on exactly this kind of model. A team I was part of was using BERT, a language model, to make sense of messy text, and there was protein folding and sequencing work going on alongside it that was some of the cleverest engineering I&amp;rsquo;ve stood near. All of it was AI by any textbook definition. None of it had a chat interface. It just did a specific, difficult, useful thing, and did it reliably.&lt;/p&gt;
&lt;h2 id="the-proof-is-in-her-camera-bag"&gt;The proof is in her camera bag
&lt;/h2&gt;&lt;p&gt;The cleanest example isn&amp;rsquo;t anything I built. It&amp;rsquo;s already in Hailey&amp;rsquo;s hands. Her Sony A7 IV does face detection, and eye detection, in the body of the camera, live, as she shoots. It has done for years. It&amp;rsquo;s so utterly ordinary that nobody thinks twice about it, nobody calls the camera &amp;ldquo;AI&amp;rdquo;, and nobody worries about where their face went. It&amp;rsquo;s a trained model welded into a physical object and taken completely for granted.&lt;/p&gt;
&lt;p&gt;The ONNX detector I just put into krites is the same species of thing: a small model that learned what an open eye looks like and gives a direct answer. The only real difference is timing. The camera does it at the instant of the shutter, and krites does it three thousand frames later, for the shots a blink spoiled that nobody caught on the day.&lt;/p&gt;
&lt;h2 id="why-the-muddle-is-worth-minding"&gt;Why the muddle is worth minding
&lt;/h2&gt;&lt;p&gt;So why does it nag at me, one slice eating the word? Two reasons.&lt;/p&gt;
&lt;p&gt;The first is that it follows the money, not the merit. The generative, chat-fronted slice is the one you can put on a slide and demo to a room, so it&amp;rsquo;s the one that gets the funding, the headlines and the name. The workhorse ML running your camera, your spam filter and your bank&amp;rsquo;s fraud checks doesn&amp;rsquo;t make a keynote, so it slides out of view, and &amp;ldquo;AI&amp;rdquo; comes to mean only the part someone&amp;rsquo;s selling. Even the firms surfacing the rest more thoughtfully are still in that lane: Google giving Imagen, Veo and Lyria their own identities is a good instinct, but those are identities for &lt;em&gt;generative&lt;/em&gt; tools. The quiet predictive stuff still doesn&amp;rsquo;t get a logo.&lt;/p&gt;
&lt;p&gt;The second reason is closer to home, and it&amp;rsquo;s about trust. When I tell Hailey that krites uses &amp;ldquo;a model&amp;rdquo; to spot blinks, the word &amp;ldquo;AI&amp;rdquo; invites her to picture four thousand irreplaceable wedding photos being shovelled into someone&amp;rsquo;s cloud chatbot. They aren&amp;rsquo;t. It&amp;rsquo;s a five-megabyte file that runs on her laptop with no network connection, finds eyes, and returns a number. When the inputs are this irreplaceable, telling those two things apart isn&amp;rsquo;t pedantry. It&amp;rsquo;s the difference between a tool she can hand the originals to and one she can&amp;rsquo;t.&lt;/p&gt;
&lt;h2 id="what-i-should-have-said"&gt;What I should have said
&lt;/h2&gt;&lt;p&gt;So, to set the record straight. The culler&amp;rsquo;s first pass is still pure arithmetic, no model, exactly as I said. But krites does machine learning now, and it&amp;rsquo;ll do more, deliberately, anywhere a learned model earns its place by giving a consistent, inspectable answer that a threshold on its own can&amp;rsquo;t.&lt;/p&gt;
&lt;p&gt;That isn&amp;rsquo;t me breaking the &amp;ldquo;no AI&amp;rdquo; promise. It&amp;rsquo;s me being precise about a word that&amp;rsquo;s stopped being precise. When I said no AI, I meant no generative AI. The machine learning was always coming. It was already in the camera that took the photo.&lt;/p&gt;</content:encoded></item></channel></rss>