Over a recording you can ask for two analyses besides the transcript: the tone — whether the voice sounds tense, cheerful, flat — and the speaker profile, which estimates characteristics such as age range or gender from the signal.
Both are useful. Both are very easily misread. And there's a line we'd rather draw ourselves before a newspaper headline draws it.
They are acoustic estimates. The system measures physical characteristics of the signal — fundamental frequency, timbre, energy, rhythm — and compares them against learned patterns.
That means two uncomfortable things:
They get things wrong. Not occasionally: perfectly routinely. A voice hoarse from a cold, an accent poorly represented in the training data, a bad phone line, someone speaking quietly because they're in an office. Any of those changes the result.
They don't measure what they appear to measure. Tone analysis doesn't say whether somebody is angry: it says their voice sounds like the voices that were labelled angry in the training data. Some people argue quietly and some people celebrate by shouting.
In aggregate and over processes, not over people:
Notice the pattern: in all three cases the result points at a process to review, not at a person to judge.
They must not be used to decide about a person. Not for hiring, not for rejecting a CV, not for scoring an employee, not for deciding who gets served first.
We don't say that out of legal caution. We say it because the system doesn't know how to do that, and using it that way produces decisions that look objective — they come out of a machine, they arrive with a number — and that are actually measuring somebody's cold, or their accent.
When an acoustic estimate enters a decision with consequences, the error stops being a percentage in a table and becomes a specific person that something happened to.
Worth stating because almost nobody warns about it: these analyses produce inferences about a person from their voice. Depending on what they're used for, they can come close to what the GDPR calls special categories of data, under article 9.
And if the result feeds an automated decision with significant effects on someone, more articles come into play. If your case comes near that line, it's a case that calls for an impact assessment before an integration.
These two features carry their warning in the documentation, on the use cases page, and in the API response itself. It isn't hidden in an appendix: it's in the same place where what they're for is explained.
We could leave it out. It would sell marginally better. But a provider that lets you build something that's going to fall over — legally, ethically or both — isn't doing you a favour; they're handing you the problem and charging you for it.
And the other half matters just as much: identifying a person by their voice is something we don't offer and won't offer. That's biometric processing, a different category of product and of risk, and almost nobody who asks for it actually needs it — what they usually need is separating speakers, which is a different thing and is available.
Anything to add or correct? Write to support@uttera.ai. If you correct us, we edit the post and credit you.