Transcribe audio
Upload a recording, get the text. It can tell the speakers apart, add timestamps and, if you ask, produce the summary and the analysis in the same pass.
One engine, two doors. A web page where the work gets done without writing any code, waiting for you the moment you sign up. And an API so your own software does it on its own, with nobody opening a browser. Below is everything it can do, one paragraph each, with a link to its manual.
One page with four tabs. Drop a file in or paste some text, press the button, download the result. It is what you will see when you walk in.
Upload a recording, get the text. It can tell the speakers apart, add timestamps and, if you ask, produce the summary and the analysis in the same pass.
Paste a text and hear it. Nine languages in the standard voice and around thirty in the high-quality one, at the speed and in the format you choose.
Upload a short sample of a voice and the text is read in that voice. Meant for dubbing yourself, not for impersonating someone else: that is forbidden, and written down.
A recording goes in one language and comes out in another. Tick the box that keeps the original voice and it comes out spoken by that same voice instead of by a stock one.
The Studio shows the estimate before you launch the job and the real figure after, in credits. No surprises at the end of the month, because there is no end of the month: it is on the same screen.
There are real limits —a three-word sentence translates worse than a fifteen-word one, and silence can make the engine invent— and they are written down instead of being discovered by surprise.
The second application, and deliberately narrow: three steps, with no options in the way. Same menu, same account.
Write the sentence you need to say tomorrow in the language you already speak, and it comes back in the one you are learning, with its phonetic transcription underneath.
The audio is generated when you ask for it, not before, at whatever speed you need. And you can play it again as often as you like without paying for it twice.
Record your attempt and it is compared sound by sound against the right one: what you did, what you should have done, and in which word. If you also want it explained in plain words, a button asks for that — and warns you it costs more before it runs.
Six services over an audio file or a text. We speak the same dialect as OpenAI's API, so in many cases changing the base URL is enough.
Audio in, text out. With speaker separation, timestamps, and the option to ask for the summary and the analysis in the same request, without uploading the file again.
Three kinds of voice, not two: standard, high quality, and cloned from a sample of your own. And you can start hearing it while it is still being generated instead of waiting for the whole file.
Text to text, or a whole recording from one language into another. And with transposition: the translation spoken by the voice of the original recording.
How much each person spoke, in what tone, at what pace.
With a limit we impose ourselves: these are acoustic estimates, they get things wrong, and they must not be used to make decisions about a person. They are there to improve a process, not to judge anyone.
An hour of audio comes back as what was agreed and what is still pending, structured so you can map it onto fields.
And you pay your LLM less along the way: 15,410 tokens of transcript against fewer than 700 of summary, measured on a real recording.
Send the sentence someone meant to say plus the audio of how they said it, and it returns which sound was made, which one was due, and in which word. It is what drives the application above, and it is open for your own language platform to use.
Sign up, create a key, send it in a header. A key can be restricted to specific IP addresses, so a leaked one is of no use from anywhere else.
There is published, tested code for Asterisk phone systems, for n8n and for AI agents. They are not toy examples: they are in production.
A free account with credits to try it out, no card. You are charged for what you consume —per second of audio, per character of text—, not per request, and every response carries a header with exactly what was billed.
Create a free account Read the documentation
Processed entirely in Spain · We do not keep your audio · We do not train on your data · Open source