Voice Comparison

Voice Comparison — Is It the Same Person Speaking?

Upload two recordings of someone talking. DiffALL builds a voiceprint from each and compares them, so it works even when the two clips are in different languages, recorded on different devices, and say completely different things.

Free, no sign-up — your files never leave the comparison.

Recording 1

Drop file here or click to browse

MP3 · WAV · FLAC · AAC · OGG · M4A · OPUS · WMA

 
Recording 2

Drop file here or click to browse

MP3 · WAV · FLAC · AAC · OGG · M4A · OPUS · WMA

 

No files handy?

Free to use — no sign-up needed

API Keys

Use the comparison API from scripts and CI. Each call costs one credit.

    Voiceprint
    Neural speaker embedding
    Any language
    Words don't have to match
    3s
    Minimum speech per clip

    What DiffALL measures

    🎙️

    Speaker Embeddings

    Each recording is reduced to a 512-dimension voiceprint by a speaker-verification network trained on thousands of speakers. Voiceprints are compared directly — no alignment, no matching words.

    🌍

    Language Independent

    The model responds to who is speaking, not what they say. Two clips in different languages compare exactly as well as two in the same one.

    🤫

    Silence Removal

    Pauses and dead air are detected and dropped before analysis, so a long silence in one file can't drag the result toward 'different'.

    ⚖️

    Honest Verdicts

    Three outcomes, not a number pretending to be a probability: likely the same speaker, likely different, or not conclusive — with the raw voiceprint similarity shown underneath.

    🚫

    Refuses Bad Input

    Under 3 seconds of detectable speech, it tells you so instead of guessing. Short clips are where speaker comparison quietly goes wrong.

    🎧

    Listen Alongside

    Both recordings play back next to the verdict, so you can check it with your own ears.

    Supported formats

    MP3 WAV FLAC AAC OGG M4A OPUS WMA

    Use cases

    Podcast Speaker Check

    Confirm which of two guests is speaking in an unlabelled clip from an old episode.

    Voice Clone Detection

    Compare a suspected synthetic clip to a known genuine recording of the same person.

    Archive Labelling

    Group unlabelled recordings in an interview archive by who is talking.

    Dataset QA

    Check that clips filed under one speaker ID in a speech dataset really do come from one speaker.

    Frequently asked questions

    How do I tell if two recordings are the same person?

    Upload both recordings and choose 'Compare the voices'. DiffALL removes silence, builds a voiceprint from the speech in each file, and compares them. You get one of three answers — likely the same speaker, likely different speakers, or not conclusive — plus the voiceprint similarity behind it.

    Do both recordings need to say the same words?

    No. This is the difference between voice comparison and audio comparison. Voice comparison measures who is speaking, so the two clips can be in different languages, on different topics, and recorded years apart.

    How much audio do I need?

    At least 3 seconds of clear speech in each file — silence and background noise don't count. Accuracy improves up to about 10 seconds. The free plan analyses 15 seconds of speech per file; Pro analyses 60.

    Is this accurate enough to use as evidence?

    No. It is a strong indicator, not proof, and it is not forensic voice identification. Accuracy drops on short clips, heavy background noise, singing, and recordings with more than one person talking. Never treat a result as a factual claim about a real person.

    Does it work on languages other than English?

    Yes. The underlying model was trained mostly on English but responds to vocal characteristics rather than words, and it separates speakers in other languages just as cleanly. We measured this on Mandarin recordings before shipping it.

    What does 'not conclusive' mean?

    The two voiceprints landed too close together to call either way. It does not mean the voices secretly match, and it does not mean they don't. The usual causes are too little speech, heavy noise, or more than one person speaking in a clip.

    Ready to compare voice?

    Free, no install, no sign-up required for your first comparison.

    Start comparing now ↑