UtteraUttera

What Uttera is actually used for

What follows isn't a list of industries invented to fill a page: it's how a system that transcribes, summarizes, translates and speaks is being used. Where data sovereignty is the thing that decides the purchase, we say so in place and explain why.

Try it without signing up Read the docs

If you work on your own

An account, a key, and the Studio page. Nothing to program.

Clinical dictation

Doctors, psychologists, physiotherapists and vets who dictate their notes after each patient and want the text, not the audio. Record on a phone, drop it into the Studio, get the transcript back.

Here sovereignty isn't a bonus feature. A clinical recording is health data — a special category under the GDPR — and running it through a service that stores it, reuses it or processes it outside the EU is a problem a generic consent form does not fix. Here it never leaves Spain and it is not stored.

Law firms and advisors

Client meetings, calls, notes from a hearing. The recording goes in and out come the transcript and a summary of what was agreed, ready to paste into the file.

Professional privilege doesn't distinguish between a filing cabinet and a server. Discarding the audio when we answer stops being a preference and becomes part of how the duty is met.

Podcasts and video

One episode yields the transcript to publish, the chapters, the subtitles and — if you want it — the same episode translated and voiced in another language.

A cloned voice starts from a sample of your own: to dub yourself into a language you don't speak, the voice doesn't have to stop being yours.

Accessibility

Read a long document aloud to listen to it while you drive, or caption a video that has none. Nine languages in the standard voice and around thirty in the high-quality one.

Studying and note-taking

A recorded two-hour lecture is transcribed in under a minute and summarized into an outline. What used to be three hours of review is ten minutes of reading.

Interviews

Journalism, qualitative research, oral history. Transcription with speaker separation, so you know who said what without listening to the whole tape again.

If you run a company

Over the API, or by connecting what you already have. There's published, tested code for everything below.

Phone support

Your phone system records, Uttera transcribes and summarizes, and the result lands where your people work. This isn't theory: there are two Asterisk AGI scripts published, one to speak to the caller and one to listen to them.

How a phone system connects →

Sales and CRM

Every call leaves a written note with the commitments and the next steps, instead of depending on someone remembering to write it. The summary comes structured, so it maps onto fields.

Quality and training

Over a sample of calls: what was said, who spoke for how long, and in what tone. To find where the script gets stuck and what needs teaching.

With a limit we impose ourselves: tone and speaker profile are acoustic estimates, they get things wrong, and they must not be used to make decisions about a person — not for hiring, not for scoring anyone. They exist to improve processes, not to judge people.

Selling abroad

Training, announcements, recorded support: made once in one language, then translated and voiced in the rest. Without booking the voice actor again every time a sentence changes.

Automating without code

If you already use n8n, there's a node and ready-made workflows to import: a folder of recordings in, a summary to email or chat out.

The code, on GitHub →

AI agents

We're OpenAI-compatible: if your agent already talks to it, you change the base URL and it can hear and speak. There's also a skill built for agents.

And one almost nobody sees: paying your LLM less

If a language model from another provider has to reason about a call, the expensive route is sending it the whole transcript. The cheap one is sending it the summary.

Measured on a real 70-minute recording: 15,410 tokens of transcript against under 700 of summary. Twenty to thirty times less. And on top of that, the recording never leaves here.

The numbers, and when it does NOT pay off →

What it costs to start

A free account with credits to try it, no card and no payments. What is charged is charged per second of audio, not per request, and the response carries in a header exactly what was billed.

See the plans →

If you're a public body

Here the question usually isn't what it can do, but where it is processed and what is kept.

Council meetings and minutes

A three-hour council session is transcribed with speaker separation and summarized by agenda item. Writing the minutes stops being an afternoon's work.

The video of the session is public, but the draft minutes are not: processing them in someone else's cloud turns an internal step into a disclosure of personal data.

Courts

Hearings, statements and appearances transcribed with who speaks at each moment. Finding one sentence inside nine hours of recording stops being an intern's job.

Healthcare

Clinical dictation and reports. With health data the decision isn't one of convenience: it is a special category under the GDPR, and where it is processed is part of the risk assessment.

Uttera transcribes and summarizes; it is not a medical device and it decides nothing clinical. What it produces is text a professional reviews.

Education

Recorded classes turned into notes and subtitles, in several languages. Accessibility for those who don't hear well and revision material for everyone, out of the same work.

Archives and heritage

Audio collections — oral history, radio, sound archives — that can't be consulted today because they aren't written down. Transcribing them makes them searchable.

Research

Interviews and focus groups transcribed and separated by speaker. Ethics boards almost always ask the same two things: where it is processed and what is kept. Both answers are written down.

Why data sovereignty decides so many of these purchases

It isn't a marketing label: it's what makes a procurement file pass or fail.

No transfer to justify

When a vendor processes personal data outside the European Economic Area, that transfer needs its own legal basis: standard contractual clauses, an impact assessment, and the uncertainty they've carried since Privacy Shield fell. Here that problem doesn't exist, because no data leaves.

Purpose limitation, for real

Article 5(1)(b) of the GDPR requires that data be processed only for what it was collected for. The industry's usual formula — "we don't use your personal data", and below it permission to use it aggregated or "anonymized" — isn't here: no copy survives to do it with.

We don't train on your data, and we don't aggregate it

Your audio and your text train no model, ours or anyone else's, and they don't turn into statistics or study material. Audio is processed in memory and discarded when we answer.

And it can be audited

The engines that do the work are published under an open license. A privacy claim you can only take on faith is worth less than one you can read.

How it holds up, article by article →

Don't see your case here?

It probably fits anyway: underneath, all of this is five services over an audio file. Write to support@uttera.ai and we'll look at it — and if it doesn't fit, we'll say that too.

Create a free account Read the docs

Processed entirely in Spain · We don't keep your audio · We don't train on your data · Open source

Complies with the GDPR and Spain's LOPDGDD