← All articles

Can a Voice Comparison Tell If Someone Is Disguising Their Voice?

People send two voice messages and ask a reasonable question: is this the same person, and are they trying to sound like someone else? A speaker comparison can answer the first part well. The second part needs an honest answer about what it measures.

What a voiceprint is built from

A speaker model turns a recording into a voiceprint — a list of numbers describing the voice, built mostly from the shape of the vocal tract: the resonances that make a voice recognisable whatever words it says. Two recordings of the same person land close together; two different people land apart.

That design is deliberately robust to how someone speaks and what they speak through, because those change all the time in real recordings.

What we measured

We took 200 pairs of recordings where both are genuinely the same person (from VoxCeleb, a public research set of thousands of speakers), altered one side of each pair, and ran both through exactly the engine the site uses.

change to one recording called “same” “not conclusive” called “different”
none 86.5% 13.5% 0%
spoken 15% faster or 10% slower 82.5–85.5% 14.5–17.5% 0%
squeezed through a phone line 80% 20% 0%
pitch shifted ±1 semitone 39.5–49.5% 49.5–58% 1–2.5%
pitch shifted ±2 semitones 2.5% 58–75.5% 22–39.5%
pitch shifted ±4 semitones 0% 13.5–19.5% 80.5–86.5%

The first three rows are the reassuring part: speed and a bad phone line almost never turn the same person into a stranger. The answer falls back to “not conclusive” when the evidence thins out, and never to a confident wrong verdict.

The last three rows are the honest part.

A voice changer moves exactly what the model listens to

The pitch shifts above work like a voice-changer app: they move pitch and the resonances together, so the voiceprint itself moves. At about two semitones the same person starts reading as someone else; at four, the comparison confidently says “different speakers” — and for that recording, the voiceprint really is different.

So a comparison cannot see through a voice changer. A “different” result on a message you suspect was disguised does not prove it was another person.

What it does not measure

A person disguising their own voice — putting on an accent, speaking higher, whispering — keeps their own vocal tract, which is a different situation from the app-style shift measured above. We have not measured it, and we will not claim a number for it.

How to read the result

  • “Same speaker” on a suspected disguise is meaningful: the disguise did not move the voice far enough to hide it.
  • “Not conclusive” is common when one recording is altered, short or noisy. Ten seconds or more of clear speech from each side is the likeliest fix; running the same files again gives the same answer.
  • “Different speakers” is reliable for ordinary recordings — on VoxCeleb’s test pairs it was confidently wrong in 9 of 33,844 decided results (0.027%) — but not for a recording you believe went through a voice changer.

Every decided result on DiffALL’s voice comparison carries a line saying this, because a confident answer to the wrong question is worse than none.

Stop hunting for differences by hand. DiffALL spots every change between any two files — automatically.

Compare your files — free