UtteraUttera

Pronounce

A page separate from the Studio, at app.uttera.ai/en/pronounce, for anyone learning a language. No formats, no voice catalogue by name, no cost estimator: someone learning German does not want to choose between opus and flac.

Uttera comes from to utter: to pronounce, to say aloud. This is that part of the name.

It solves a problem learning apps do not touch: they only let you hear their own content. You can repeat their lesson, but you cannot ask them how the sentence you need to say tomorrow sounds. Here you write that sentence.

The three steps

StepWhat happens
1. Write it in your own languageThe page is in your language and you write in yours. You pick which one you want to learn to say it in. Up to 300 characters: a sentence, not a text.
2. This is how you say itThe sentence appears in the language you are learning together with its phonetic transcription. That is where you stop and read. The audio is generated separately, when you want it and at the speed you choose: reading the sentence is free, generating it is not, and paying for audio of something you have not read yet makes no sense.
3. Now say it yourselfYou record yourself — browser microphone or file — and we give you back what we understood, marking word by word what came out and what did not.
What the third step gives you is a mirror, not a grade. If you wrote “rojo” and the transcript says “rojo”, you said it recognisably. If it says “rollo”, you now know which sound is failing you. There is no score, and that is deliberate: for learning, an honest mirror beats a number.

What it costs

There is no estimate beforehand, on purpose: someone practising is not evaluating a service, and a figure before every sentence would be noise. After each operation we say what it actually cost. A full cycle — translate, listen and check yourself — is around 0.17 credits for an ordinary sentence.

The limits, said up front

“Why doesn't it understand me?”: the AI analysis

The three steps give you a mirror: which words came out and which did not. When that is not enough — you know something is off but not what — a separate button compares sound by sound what you said against what it should be, and explains it.

It does not use ordinary speech recognition, and the reason matters: that is built to understand you, not to measure you. It corrects your pronunciation towards the word it thinks you meant, which is exactly the error we want to see here. The analysis uses a phoneme recogniser, which writes what it hears even when it is not a word, and compares that string with the one that should have sounded.

What you get back: the percentage of sounds that came out right, the list of those that did not — you made [s] where it should be [θ], and in which words —, the written explanation, and both phonetic transcriptions for anyone who can read them.

It costs considerably more than the rest, which is why you ask for it rather than getting it automatically. The word mirror costs cents of a credit; this is around 3, because your recording has to be phonemised and the explanation drafted with a language model. You are told before it runs, and you can decline. The exact charge appears afterwards, as with everything else.

The idea is that you practise many times for almost nothing and ask for the analysis when you want to know what you are doing wrong, not on every repeat.

If the recording does not sound like the sentence, it is not analysed. This happens more than you would think: the wrong microphone picking up the television, or the system mix. In that case we tell you and the analysis is not charged, rather than explaining sound by sound a mismatch that means nothing.
What this is NOT.

It is not a grade or an exam: it is a comparison between two strings of sounds. And the explanation is drafted by a language model from that comparison — it does not decide what is wrong, that is computed beforehand — so read it for what it is: help in understanding the data, not a verdict.

It does not replace a teacher either. For the sound of an ordinary sentence it is more than enough; to polish one exam vowel, a teacher is still a teacher.

What you upload is transcribed and deleted: recording yourself practising is exactly the kind of audio that should not sit in anyone's archive — Privacy.