Like all the best jokes, it’s actually true – computers don’t like a Scottish accent.
As a nation, it’s a problem that unites us and enrages us. Who hasn’t chuckled to themselves in an empty lift, whispering “Elevuhn!” to a bot that isn’t there? And who doesn’t dread hearing “Sorry, I didn’t catch that” when calling a government hotline?
But voice tech has come a long way since Burnistoun, and we Scots no longer have to live in fear of spelling out our postcode.
Like those with strong regional accents up and down the country, developments in speech recognition coupled with AI have levelled the playing field – now we’re easily mimicked, rather than misunderstood.
Unfortunately, that’s opened the door to a new set of problems.
Aye, Robot
“You don’t expect technology to sound so authentic,” said Dr Neil Kirk, chartered psychologist and reader in psychology at Abertay University.
“If technology has never spoken like you, if technology has never been able to understand you, why would you expect it to speak like you?”
This is the central tenet behind Kirk’s recent work, which examines the subtle ways our brains react when AI speaks with a familiar voice. His research has found that when people hear a voice that sounds like them, they are far more likely to assume it’s human – even when it’s been artificially enhanced.
It’s the sharp edge where emerging technology meets our squishy, primordial instincts, but this ‘default to human’ bias, Kirk warns, isn’t trivial: it’s a vulnerability waiting to be weaponised.
Following experiments where listeners were given different voice recordings, some human and some transformed using AI, Kirk found that people struggle to accurately distinguish synthetic voices.
In other words, our brains are wired to trust the familiar, even when the familiar is artificially constructed.
“People couldn’t tell the original from the new AI-transformed voice, but they just kept assuming that they were original human voices because they sounded so good,” said Kirk, adding that the effect was even stronger when the voices spoke with familiar dialects (in this case, Dundonian Scots).
“It’s not what they’re hearing, it’s what they’re believing, and that is that it sounds really local, and so that has to be a real voice.”
Through the Uncanny Valley
For security professionals, the implications are obvious, and, to be honest, already well known. Vishing isn’t a new phenomenon, and humans have long been considered the weak link in cyber.
Last year, a global survey by McAfee found that out of 7,000 people, 70% weren’t confident they could tell the difference between a cloned voice and the real thing – a scary stat, particularly when the same study also cites that just three seconds of audio was found to be enough to produce a clone with an 85% voice match to an original source.
But while AI-enabled speech models can create realistic impersonations from just a few seconds of target audio, resulting in voice clones that sound like the panicked cries of friends and family or even the President, telling citizens to skip voting, those instances are highly targeted.
What Kirk has discovered is a more subversive danger.
“If I hear a Southern British English accent, I might immediately think that sounds like a scam. Whereas if I hear somebody who sounds a lot more local to me, I might think it’s actually somebody from my local branch of the bank,” he explains.
“The vulnerability comes from bypassing people’s scepticism.”
Kirk’s research is among the first to recognise this emerging threat. His findings suggest that a level of authenticity can be achieved with a free subscription to ElevenLabs that might be enough to get past the psychological defences of frontline staff and customers, just by tweaking an AI-enhanced voice to sound like it comes from the same local area.
“We hear a lot about AI voice scams and fraud that involve voice cloning of actual individuals, like targeted scams using a relative’s voice or maybe a high-profile person in a company,” said Kirk, “but that type of scam is more resource-intensive. You need to really train on a specific voice and a specific identity.
“Whereas voice cloning scams are very tailored, my feeling is this could be done on a much bigger scale, with a much lower barrier to entry.”
But how much bigger could the problem become? Research from Starling Bank provides some indication. Last year, the challenger bank found that although more than a quarter of UK adults have been targeted by an AI voice cloning scam, almost half (46%) didn’t even know this type of scam exists.
While Starling catastrophises this into potentially millions of us being at risk of falling victim to AI voice cloning, Kirk’s study proves that there is an as-yet untapped seam of criminality where fraudsters use our own identities against us.
“I’ve stumbled on something that maybe hasn’t been utilised in a scam, and I need to be careful that I’m not actually promoting a how-to guide,” he said, “but hopefully, we can stay one step ahead in mitigating the problem.”
Nudge Against the Machine
There is a slew of articles and advice from every corner of the cyber and tech industry that claims to have the secret to detecting deepfakes, and plenty of firms offering security awareness training to defend against the use of synthetic voices and images.
But how effective is it? According to Kirk, not very.
“The reality,” he wrote in his paper, “is that AI-generated voices are now becoming so advanced that it may be nearly impossible, not to mention costly, to train people to reliably distinguish them from real human speech.”
That sentiment is reiterated in Kirk’s latest research, which claims that ‘training and expertise offer little protection against AI voice deception’, citing studies that found a measly 3% improvement in AI detection when people are taught to look for deepfakes, even when using examples.
Recommended reading
- How Worried Are Brits About AI-Fuelled Phishing?
- ‘Stealth’ Phishing Attacks Rise As Over Half Evade Detection
- Report: More Than 50% of Fraud Is Now Driven by AI
Luckily for us all, Kirk has already been hard at work identifying the beginnings of a possible solution, and it’s strikingly simple.
Based on his research, again using AI-enhanced Dundonian Scots accents, Kirk reckons that a more effective strategy for tackling deepfakes, voice-based or otherwise, may be to subtly shift people’s default assumptions.
Using ‘nudges’, basically small hints or cues reminding people that AI can sound naturally human, Kirk managed to pretty much double instances of synthetic voice recognition.
Rather than trying to train people how to spot a fake voice, Kirk simply let his participants know that AI is now sophisticated enough to mimic our own accents, challenging pre-existing assumptions that, in his words, ‘AI sounds robotic, American, and quite stilted.’
“For me, we have relatively low-hanging fruit here,” said Kirk. “Seeing how effective these simple nudges have been is encouraging, but we’ve still got a long way to go to let people know what AI voice technology can actually sound like.
“You have to update their belief that it can speak with an accent and dialect. It does change their behaviour, and makes them less likely to trust that a voice is really human – whereas warning them to be careful does nothing.”
It’s not a technical problem to solve with more AI tools or inflated budgets; it’s psychological. The capabilities of AI have overtaken us with ferocious speed, something like if internal combustion had been invented days after the wheel, and from an evolutionary perspective, we can’t keep up.
Looks like it’s our turn to be rewired.





