Recognition comes before verification
Auditory memory is built for efficiency. Your brain matches tone, pace, and distinctive pauses to a stored pattern and answers "that's Mom" before you've had time to question it. That's not gullibility - it's how recognition normally works: in everyday life, a voice almost never lies, and checking every time would be too costly.
The problem is that recognition activates a whole package of expectations. If the voice is familiar, then the intentions must be familiar too, the context must be clear, and the request must be appropriate. You're no longer responding to a sound, but to your whole history with that person. That's the package scammers use.
Only that the sound matches a pattern in your memory.
That it is that person, they are okay, calling of their own free will, and telling the truth.
The gap between those two columns used to be tiny: faking a voice was difficult. Now a short recording from a public profile is enough - and the gap has widened. Psychology hasn't changed; the cost of faking has.
What works next is not the voice, but the rush
A familiar tone does not push you to do anything on its own. It only removes the first barrier. Then comes a structure that is nearly the same in every scenario: urgency that leaves no time to think, secrecy that leaves no one to call, and a request that must be carried out right now.
Each element targets one ability - the ability to stop and check. Urgency takes away time. Secrecy cuts you off from a second opinion. And a recognizable voice makes the very idea of checking feel awkward: who calls Mom back to make sure it's Mom?
- Recognition
- Urgency
- Secrecy
- Request
If all four come together in a conversation, that is a sign in itself, no matter whose voice you hear. Someone close to you can handle a two-minute pause. A pressure script cannot, which is why it won't give you one.
How well established is this?
Experiments have measured how poorly people distinguish synthesized speech from real speech: participants guessed better than chance, but were wrong in about every fourth case, and a short training session helped only a little. This is the lab, not a panicked call: short recordings, a couple of languages, no pressure, and no familiar history behind the voice.
Hearing is unreliable at spotting fakes: there are still few measurements, all under laboratory conditions
That does not change the direction of the conclusion - if anything, the opposite: in a real conversation, with urgency and anxiety, you will not listen more carefully than in a lab. It is not wise to make your own hearing the expert on authenticity.
Finish reading in the Miqo app
1 more minute, then a quiz at the end. Find it under “Psychology → Social Engineering”
Point your phone camera here to open the App Store or Google Play.
Sources
- Mai K. T. et al., "Warning: Humans cannot reliably detect speech deepfakes," PLOS ONE, 2023
- on principles of influence - R. Cialdini, "Influence."