Audio to Text — transcription in paragraphs with speaker labels

An online tool that transcribes audio and video recordings — meetings, interviews, lectures, podcasts and voice notes — into text. The result comes in paragraphs with speaker labels and timecodes; you can edit it on the page, copy it, or download it as DOCX, PDF, TXT or SRT/VTT subtitles. Short files are recognized in the browser; hours-long ones run on our server and you get an e-mail when the text is ready.

How to transcribe audio to text

1
Upload the recording

MP3, WAV, M4A, FLAC, OGG, WebM or a video — drag the file onto the page or pick it on your device.

2
Set up the recognition

Choose the language and quality, tick "Label speakers" and enter the number of participants if you know it.

3
Take the text

Fix words and speaker names right on the page, then copy the text or download DOCX, PDF, TXT, SRT or VTT.

Turn a recorded conversation into a readable document: who said what, and when

Want to dictate live instead? 🎤 Type by voice into the page

Drop audio file here
MP4, WEBM, MOV, MP3, WAV, M4A, OGG, FLAC, AAC, M4B, WMA, AIFF, OPUS, CAF, MKV, AVI, WMV, FLV, M4V, 3GP, TS, MTS, M2TS, VOB, MPG, MPEG, OGV

Your transcribed text will appear here.

00:00

Which model to pick, and what it costs you

Recognition runs in your browser, so the model is downloaded once and cached. Bigger models hear accents and noise better but take longer on the same machine.

Model One-time download Time per minute of audio When to pick it
Fast ~75 MB about 1 min Quick draft, clear speech, one speaker
Accurate ~150 MB about 3 min Everyday recordings — the usual compromise
High Accuracy (small) ~250 MB about 6 min Accents, background noise, several speakers
Premium (Large v3 Turbo) ~600 MB depends on the GPU Hardest audio; needs a GPU or a modern CPU

A video file is stripped to its audio track before anything is processed: an hour of phone video is 2–3 GB, of which the sound is about 30 MB.

Published Updated Author: