Voice Comparison — Is It the Same Person Speaking?
Upload two recordings of someone talking. DiffALL builds a voiceprint from each and compares them, so it works even when the two clips are in different languages, recorded on different devices, and say completely different things.
Free, no sign-up — your files never leave the comparison.
Drop file here or click to browse
MP3 · WAV · FLAC · AAC · OGG · M4A · OPUS · WMA
Drop file here or click to browse
MP3 · WAV · FLAC · AAC · OGG · M4A · OPUS · WMA
Free to use — no sign-up needed
Processing...
API Keys
Use the comparison API from scripts and CI. Each call costs one credit.
Copy this key now — for security it won't be shown again.
No keys yet — create one to start using the API.
What DiffALL measures
Speaker Embeddings
Each recording is reduced to a 512-dimension voiceprint by a speaker-verification network trained on thousands of speakers. Voiceprints are compared directly — no alignment, no matching words.
Language Independent
The model responds to who is speaking, not what they say. Two clips in different languages compare exactly as well as two in the same one.
Silence Removal
Pauses and dead air are detected and dropped before analysis, so a long silence in one file can't drag the result toward 'different'.
Honest Verdicts
Three outcomes, not a number pretending to be a probability: likely the same speaker, likely different, or not conclusive — with the raw voiceprint similarity shown underneath.
Refuses Bad Input
Under 3 seconds of detectable speech, it tells you so instead of guessing. Short clips are where speaker comparison quietly goes wrong.
Listen Alongside
Both recordings play back next to the verdict, so you can check it with your own ears.
Supported formats
Use cases
Podcast Speaker Check
Confirm which of two guests is speaking in an unlabelled clip from an old episode.
Voice Clone Detection
Compare a suspected synthetic clip to a known genuine recording of the same person.
Archive Labelling
Group unlabelled recordings in an interview archive by who is talking.
Dataset QA
Check that clips filed under one speaker ID in a speech dataset really do come from one speaker.
Frequently asked questions
How do I tell if two recordings are the same person?
Upload both recordings and choose 'Compare the voices'. DiffALL removes silence, builds a voiceprint from the speech in each file, and compares them. You get one of three answers — likely the same speaker, likely different speakers, or not conclusive — plus the voiceprint similarity behind it.
Do both recordings need to say the same words?
No. This is the difference between voice comparison and audio comparison. Voice comparison measures who is speaking, so the two clips can be in different languages, on different topics, and recorded years apart.
How much audio do I need?
At least 3 seconds of clear speech in each file — silence and background noise don't count. Accuracy improves up to about 10 seconds. The free plan analyses 15 seconds of speech per file; Pro analyses 60.
Is this accurate enough to use as evidence?
No. It is a strong indicator, not proof, and it is not forensic voice identification. Accuracy drops on short clips, heavy background noise, singing, and recordings with more than one person talking. Never treat a result as a factual claim about a real person.
Does it work on languages other than English?
Yes. The underlying model was trained mostly on English but responds to vocal characteristics rather than words, and it separates speakers in other languages just as cleanly. We measured this on Mandarin recordings before shipping it.
What does 'not conclusive' mean?
The two voiceprints landed too close together to call either way. It does not mean the voices secretly match, and it does not mean they don't. The usual causes are too little speech, heavy noise, or more than one person speaking in a clip.
Ready to compare voice?
Free, no install, no sign-up required for your first comparison.
Start comparing now ↑