Speech to Text

An interview, a lecture, a voice message you would rather read than replay. The recording stays on this device — the model comes to it, not the other way round. That costs about 68 MB on the first run, and the number is written on the button, not hidden behind it.

Audio or video — the browser will take the sound track itself.

The recording is not uploaded anywhere: the model is downloaded to your browser and does the listening there. After the first run it works with the network off.

Nearby: Voice Recorder · Video to MP3 · Audio Converter · Online Notepad

How to use

  1. Choose the recording — audio or video, the sound track will be taken out for you.
  2. Pick the language spoken. On the first run the model is downloaded once.
  3. Press the button, read the transcript and download it as text or subtitles.

Good to know

What it hears well and what it does not

On clean English speech the light model misses about one word in twenty — a transcript you can read and fix in a minute. On Russian, Japanese or Korean it misses about a third of the words even on studio-clean input, because the light model is the only one that fits in a browser at all. The page says which case you are in before you spend the download, not after.

Why 68 MB and no server

Every service that transcribes “instantly” sends your recording to someone else’s machine. Interviews, medical appointments, family voice messages — that is a lot to hand over for a convenience. Here the model comes to the file instead: one download, then it works offline, and the recording never becomes anyone’s dataset. The price of that choice is the download, and it is stated up front.

Frequently asked questions

Is my recording uploaded?

No. The model is downloaded to your browser and runs there. After the first run the page works with the network off.

Why is the Russian transcript so rough?

Because the light model is weak at it: about a third of the words come out wrong even on clean speech. Stronger models exist, but they are four times heavier and will not load in a browser. The page warns you before you download anything.

Can I get subtitles?

Yes — an SRT file with timestamps comes out along with the plain text. Timing is per thirty-second window, which is what the model actually knows.

How long a recording can it take?

Up to an hour. Longer files do not fit in browser memory in one piece; split them first.

Related tools