What Is Speaker Diarisation (and Why Does It Matter)?
If you've ever received a transcript that looks like one unbroken wall of text, you know how hard it is to follow who said what. Speaker diarisation solves this by automatically detecting speaker changes and labelling each segment: Speaker A, Speaker B, and so on.
How it works
Under the hood, diarisation uses a combination of acoustic modelling and clustering. The system analyses the audio signal to identify unique voice characteristics — pitch patterns, rhythm, formant frequencies — and groups segments with similar characteristics together. It doesn't know who the speakers are, only that two segments were spoken by the same person.
What it's good at
- Podcast interviews with 2–4 distinct speakers.
- Research focus groups where speakers take clear turns.
- One-to-one interviews and sales call recordings.
Where it struggles
- Overlapping speech. If two people talk at once, the system often assigns the segment to the wrong speaker.
- Very similar voices. Identical twins, or a presenter whose voice changes significantly between segments.
- Large groups. Accuracy drops off when there are more than six speakers.
- Short utterances. A single word or a brief "mm-hmm" doesn't give the model enough signal to assign confidently.
Using diarisation on Scriptring
When you upload a file via AI Transcribe, speaker diarisation is enabled by default. The output groups utterances by speaker and colour-codes them in the viewer. You can click any timestamp to jump to that point in the audio, which makes correction fast — simply listen to the segment and reassign if the label is wrong.
For critical recordings, our human transcription service applies manual diarisation with consistent speaker label conventions, including the option to use real names rather than generic labels.
Ready to try it yourself?
Free AI transcription — no account needed. Human review from $1.25/min.