UtteraUttera

What "we don't train models" actually means

September 15, 2026

Our home page carried a line of badges:

Processed entirely in Spain · We don't keep your audio · We don't train models · Open source

It sounds good. And it's wrong.

The problem with the short sentence

"We don't train models", flat out, asserts two things at once, and only one of them is the thing we want to assert:

  1. That our customers' data never enters the training of any model. That is true and it is the real commitment.
  2. That we will never train anything. That is neither true nor something we want to be true: training a model is legitimate, it is how this technology advances, and one day we will want to fine-tune one so that Spanish is pronounced better.

A marketing sentence that closes a reasonable technical door on you is a badly written sentence. And one you may have to take back is worse than saying nothing.

What we did: go through all five

We looked for every place where the site asserted something about training. Five. Of those, four were already bounded by their subject — "your audio and your text train no model", "data is processed to give you the service and for nothing else" — and were fine.

The fifth, the one in the badge line, was the only unbounded one. And it was precisely the worst, because a list of four badges separated by dots is what gets quoted out of context, copied into a screenshot, and ends up in a tweet.

It now says "We don't train on your data". Same length, same rhythm, and it asserts more.

And a sentence that was going to become false

Along the way another one turned up, in two places:

The models come pre-trained and learn nothing from what passes through them.

The first half stops being true the day we fine-tune one of our own. An expired sentence on a privacy page is worse than not having it, because nobody reads it again and it stays there.

We changed it for what holds up whatever happens, and is what actually matters to the customer: the models learn nothing from what passes through them. Every request begins and ends without leaving a trace in the model, whoever trained it and with whatever.

The honest version is the stronger one

The documentation now has a section called "What this does not say", and it says this:

That distinction — between your data and licensed data — is precisely the one almost nobody makes, and it is more credible than the absolute. Nobody believes the absolute; this one can be defended.

A note about the industry

The usual formula is "we don't use your personal data" and, underneath, permission to use that same data once aggregated or "anonymized". An audio file with the customer's name stripped is still a person's voice. If that door is open, the promise above doesn't mean much.

We would rather the promise didn't depend on our good intentions: audio is processed in memory and discarded when we answer, so no copy survives to do it with. That isn't a policy, it's a consequence — and consequences don't change with a terms-update notice.

Anything to add or correct? Write to support@uttera.ai. If you correct us, we edit the post and credit you.

← All posts