UtteraUttera

Diarization tells you how many are speaking and when, not who they are

September 15, 2026

When you ask for a transcript with speaker separation, what you get back looks like this:

S0 — Good morning, I'm calling about order 4471. S1 — One moment, let me check. S0 — No rush.

S0 and S1 aren't names. They're labels. And that's where all the confusion lives — the confusion worth clearing up before somebody builds something on top of it.

What diarization does do

It detects how many distinct voices are in the recording and when each one speaks. It's a clustering problem: the system separates stretches of audio by voice characteristics and decides which ones belong to the same speaker.

That's already worth a lot. It turns a wall of text into a readable dialogue, it lets you measure who spoke for how long, and it makes finding one sentence inside nine hours of recording a possible task.

What it doesn't do

It doesn't know who S0 is. It has no idea. There is no database of voices behind it to compare against, and there isn't going to be: that would be a biometric register, which is a different category of product and a different category of legal problem.

It also doesn't guarantee that S0 is the same person across two different recordings. The labels are local to each audio file. In Tuesday's call S0 might be the customer and in Wednesday's the agent.

And it isn't infallible within a single recording either: two similar voices can merge into one label, and a single voice that shifts register a lot can split into two.

How it becomes "who" without inventing anything

Almost always, with context you already have and the audio doesn't:

It's less elegant than asking the model for the name, but it has an enormous advantage: it doesn't make it up.

The line we don't cross

It might look as though the natural next step is identifying speakers by their voice. It isn't, and we aren't going to offer it: identifying a person by a physical characteristic is biometric data processing, which the GDPR puts in special categories and which demands a completely different regime — and which, above all, almost nobody who asks for it actually needs.

The same applies to what we do offer and should be used with care: the speaker profile and the tone analysis are acoustic estimates, they get things wrong, and they must not be used to decide anything about a person. It's written in the documentation and we repeat it everywhere the feature appears.

Anything to add or correct? Write to support@uttera.ai. If you correct us, we edit the post and credit you.

← All posts