Trust on the line…

Downton Abbey star Hugh Bonneville would quite like to keep his voice.

So would Nicola Coughlan, Siobhán McSweeney, Matt Lucas and dozens of other performers currently backing the UK-based Save Our Voices Now. This campaign is calling for stronger protection against the unauthorised cloning of people’s voices using AI.

At one level, this is a very individual problem. A voice is unusually personal. We use it to recognise people we know, decide who we trust, and work out who’s speaking even when we can’t see them. But for performers, it’s also a professional problem. Their voice can be their livelihood. And, of course, it can now be copied.

Across the UK we’re seeing a range of local and national initiatives to tackle this issue. On 8 September, Lancashire Constabulary launched Fake or Real? Know the Deal, a public-awareness campaign designed to help young people and families recognise and respond to AI-generated deepfakes. Importantly for us, Lancashire’s definition does not stop at manipulated photographs or video. It explicitly includes audio: clips made to sound as though somebody said something they never actually said.

The force has also published a Fake or Real? toolkit because being safe online is not about magically spotting every fake. It’s about the ordinary person – child or adult – knowing when to question what they’re seeing or hearing, checking sources, and understanding what to do when something feels wrong.

Climb one more rung, and the same problem becomes national. The Home Office Deepfake Detection Challenge has brought together government, law enforcement, national-security users, academics and technology companies to test how well current tools can distinguish real from synthetic media, including audio, under realistic, time-sensitive conditions.

That national work matters because this problem keeps evolving so swiftly. A recent Department for Science, Innovation and Technology review describes deepfake detection as still being at an early stage, with rapidly changing generation techniques continuing to challenge the tools intended to identify them.

So, from one person asking who owns their voice, to a local police force asking people to think harder about what is real, to the Home Office testing national detection capability, we keep arriving at the same awkward question:

What exactly are we trying to detect?

As we note elsewhere on this blog, cloning a voice is one thing. Making it say a sentence is another. But human conversation takes us into a third and arguably far more complex dimension.

Humans hesitate. We interrupt each other. We talk over one another. Mishear things. Repair ourselves halfway through a sentence. Repeat ourselves. Speed up. Slow down. Laugh. Breathe. Stall for time. Change direction. Repeat ourselves. Forget the other four things we were going to add into this list. And somehow coordinate all of this with another person in milliseconds.

So what happens when we ask AI to reproduce that?

That’s the problem at the heart of HackaCon. This isn’t just a task where you recreate the voices of Agent Luke and Chris Nemesis. That’s arguably the easy part. The real trick is making them have a convincing, spontaneous-sounding conversation that never actually happened.

Can you make the interaction feel natural? Can you reproduce the tiny timing decisions and irregularities that make two people sound as though they are genuinely responding to each other rather than taking turns reading some 1990s sitcom dialogue? And, perhaps most interestingly: what tells will give you away?

There are now just over two weeks left to find out. The HackaCon submission window closes at 23:59 UTC on Wednesday 30 September.

You don’t need a finished system before you begin. Try something. Break it. Listen closely. Change the timing. Add awkward bits. Edit them out. Discover which parts are surprisingly easy and which stubbornly refuse to sound human.

In other words: get playing, and good luck!