What we are learning as we build this. No press releases: bugs that cost us a day, measurements that came out the opposite of what we expected, and decisions that need explaining.
Give it empty audio and it doesn't return an empty string: it invents a sentence. And the worst part isn't paying for it, it's your automation acting on it.
It started as a tool for internal use in a Spanish family business group. It wasn't designed to be sold. That's why it's built the way it is.
Billing is per second of audio, not per request. That changes how it pays to split the work, and there's a trick that saves you money before you send anything.
The difference between "separating speakers" and "identifying people" takes two paragraphs to explain and saves an expensive disappointment.
The short sentence said more than we meant, and it closed a legitimate door on us. We went through all five places it appeared on the site.
Your phone system is already writing .wav files somewhere. Pointing a script at that folder is half an hour. Putting a voice inside the call is a day, and here are the five traps.
They are acoustic estimates. They get things wrong. And there's one use we won't support even if someone pays for it.
We keep the audio we generate for one hour, and it costs 10%. An hour of cache is not an hour of access to your data, and you can switch it off on every request.
A 70-minute recording is 15,410 tokens of transcript and under 700 of summary. Twenty to thirty times less, and the recording never leaves here.
The engines that move Uttera are open source. This is what it takes to stand them up on your own machine, without going through us.
It isn't a quality decision. It's about how each one uses GPU memory, and choosing wrong is the fastest way to lose an afternoon.
Four scripts and an instructions file. The hard part isn't the API: it's telling the agent when to use each thing and what not to believe.
Same routes, same parameters, same voice names. What you change is the base URL, the key… and one more thing that isn't obvious.
It isn't a marketing label. It's the difference between a procurement file that passes and one that sits waiting for a legal opinion.
The technical part is solved with ten seconds of clean audio. The part almost nobody writes down is the other one: that a voice is an attribute of a person.