There’s an ambulance going past. There was a police car about four minutes ago, and there’ll be another one along shortly, because that is simply what this road does.
Which matters more than it should, because a while back I spent an afternoon sat in a wardrobe. In a headset. In the dark. Reading sentences off a phone screen so a machine could learn to talk like me.
Not the glamorous end of software, that.
Why there’s a clone at all
Two reasons, and I’d rather put them up front than have them read as excuses at the bottom.
The first is scheduling and a room. I’ve been playing with social media (badly) and creating reels & shorts for various platforms, and most of them want some form of narration. I’m frequently not around to record it, and more to the point I’ve nowhere sensible to record it in. No treated room, no booth, nothing you’d call soundproof. Just a house on a road that emergency vehicles are extremely fond of. Hence the wardrobe… which is, genuinely, the best acoustic environment I own: soft, small, full of coats, and mercifully further from the window.
The second is that I need to know how good this stuff actually is. Not in the abstract, and not by reading what the vendor says about itself. I’m building things that might one day want a synthetic voice in them, and “is this good enough to put in front of a real person?” is a question I’d sooner answer with my own ears and my own money, before it turns into somebody else’s problem.
To be clear about one of them, since I announced it four days ago: Scout never talks to players. Ever! That’s a design position I’m not walking back, and no amount of good voice synthesis is going to change it.
It may well end up talking to the DM, though, if my designs for it go the way I intend. Whispering in one person’s ear is rather the whole job. And there’s a sister service further down the line that might want a voice of its own for a completely different reason. Either way I’d rather know where the ceiling is before I get there than find out afterwards.
So: not a vanity project. A workaround for a room, and a piece of homework.
What “wrong” sounds like
Here’s the thing nobody warns you about. Like most people, I hate the sound of my own voice. Sets my teeth right on edge. I’ve long since made peace with that.
What I hadn’t made peace with was hearing my own voice wrong.
Sometimes it comes back American. Not subtly, either. Sometimes it’s my speech pattern, my phrasing, my rhythm… pitched about two octaves too high, which is a genuinely strange thing to sit and listen to.
But the one that gets me, the one that made me stop and rewind, is this: I have a lisp.
It’s not prominent. Most people never notice it, and I’d be surprised if you’d clock it in a pub. It’s an ever so small thickness on my sibilants, a slight weight where there shouldn’t be one. It’s been there my whole life. And every so often the clone just… doesn’t do it. Renders the sentence perfectly, cleanly, crisply, without it.
And I sound like an axe murderer.
I can’t put it any better than that. Something about the sibilants landing that clean turns me into somebody you would not want to be alone with. It’s my voice, saying my words, with one tiny thing missing, and what comes out the other end is a stranger.
It isn’t the uncanny valley thing
I assumed, before anyone asked me, that what bothered me here was something big and philosophical. Hearing a machine be me. The vertigo of it.
It isn’t. I’ve thought about it properly and it’s much more boring: it’s accuracy.
I put real work into getting it right. The wardrobe, the headset, the retakes, the hand-built phonetic spellings for all the words I say constantly and it mangles constantly. So when it wanders off, it isn’t existential dread. It’s a picture hung slightly crooked. My own personal variant of OCD, which demands the thing be just right, and will not let it go until it is.
That’s a much less impressive answer than the one I expected to give. It’s also the true one.
Where I draw the line
I’m not precious about it. I am fairly specific, though.
Short-form is fine. Soundbites, a reel, thirty seconds over a cartoon avatar that is very obviously a drawing and isn’t pretending to be a photograph of a man. Those are still my words. That’s still the meaning I set out to convey, and nobody watching is being told a lie about what they’re looking at. (Getting a clone to carry actual emotion is hard work, mind, and you can usually hear it trying.)
Long form isn’t fine. Anything that runs on a bit, anything meant to portray real life… that has to actually be me. Sirens and all. I’m not out to bamboozle anybody. And frankly, no voice clone can yet carry me and all my idiosyncrasies as well as I can, so on the pieces where those idiosyncrasies are the point, using one would be a downgrade as well as a fib.
Note the “yet”. I’m not about to stand here and tell you humans have got some permanent edge on this. We probably haven’t. But we have today, and today is when I’m publishing.
Why it matters that only I can hear it
Which brings me to the bit that took me a while to work out, because on the face of it none of the above should matter at all.
Nobody else can hear the lisp. Nobody is listening to a reel going “hmm, sibilants seem light this week”. If I let it drift, precisely zero people would write in. So who exactly am I doing this for?
I did my time on the speaker circuit. Never packed a room out in my life, but I stood up in front of people for years, and there are folk out there who know me and know what I sound like. And there will be more of them: people who find a reel before they ever find me, and who then, at some conference or meetup or pub, actually meet the bloke.
When that happens, the two need to match.
That’s all this is, really. The clone isn’t standing in for my identity, it’s standing in for me, at a particular moment, and it’s making a small promise on my behalf about what I’ll be like when you meet me. If the recording has a lisp and the man doesn’t, or the recording is American and the man is thoroughly not, then the promise was rubbish and you’ve every right to feel slightly conned.
Fidelity matters. No continuity errors allowed.
Which is, I notice, exactly the language I’d use at a game table. Continuity errors are the thing that breaks a campaign: not the big dramatic mistakes, but the small ones nobody flags at the time and everybody quietly stops trusting you over. Get the barkeep’s name wrong twice and the tavern stops being a place.
Same job here. It’s my voice. It should keep saying the same thing about me whether I’m in the room or not.
Even if I’m the only one holding the continuity notes.





