Speaker diarization

Speaker diarization: how software knows who spoke when

Speaker diarization is the process of splitting a recording by speaker. In simple terms, transcription answers “what was said?” while diarization answers “who spoke when?”

Transcription and diarization solve different problems

Speech-to-text turns sound into words. Diarization groups those words into Speaker 1, Speaker 2 and so on. Combining both turns a long transcript into something that reads like a conversation.

Speaker recognition goes one step further

Diarization can tell that two voices are different without knowing their names. Speaker recognition links a voice to a known person, such as recognizing that Speaker 2 is Alice because you named her in an earlier recording.

Why it matters for meetings and interviews

When decisions, quotes and action items depend on who said them, a plain transcript is incomplete. Speaker labels make it faster to verify statements, review interviews and follow a multi-person discussion.

Diarization is not perfect

Overlapping speech, distant microphones, background noise and very similar voices can cause mistakes. A good app should let you rename, merge and reassign speakers instead of treating the first result as final.

Start with your next conversation.

Download Murmur, download a speech model once, and every recording after that is transcribed on your own device — meetings, interviews, lectures, ideas on a walk.

Free to download · No account · Works with no signal

On iPhone and iPad the free version transcribes the first 20 minutes of each recording; Murmur Pro is a one-time purchase that lifts the limit. On Android, transcription is currently free with no recording-length limit.