Audio to Text — transcription in paragraphs with speaker labels
An online tool that transcribes audio and video recordings — meetings, interviews, lectures, podcasts and voice notes — into text. The result comes in paragraphs with speaker labels and timecodes; you can edit it on the page, copy it, or download it as DOCX, PDF, TXT or SRT/VTT subtitles. Short files are recognized in the browser; hours-long ones run on our server and you get an e-mail when the text is ready.
How to transcribe audio to text
MP3, WAV, M4A, FLAC, OGG, WebM or a video — drag the file onto the page or pick it on your device.
Choose the language and quality, tick "Label speakers" and enter the number of participants if you know it.
Fix words and speaker names right on the page, then copy the text or download DOCX, PDF, TXT, SRT or VTT.
Turn a recorded conversation into a readable document: who said what, and when
Want to dictate live instead? 🎤 Type by voice into the page
Files are transcribed one by one; every finished text is saved to the history below. Each file is a separate run.
📝 Names & terms (optional)
Your transcribed text will appear here.
Transcription
The protocol is written by a language model. For the analysis the text goes to a third-party AI provider (Anthropic, USA). Before sending we replace passport numbers, tax and insurance IDs, account and card numbers, phones and e-mails with placeholders; names, company names and addresses stay in the text — do not upload a conversation you are not entitled to share with third parties. Every protocol item is checked: it is shown only if its quote appears verbatim in your text. That rules out invented conclusions, but it does not make the analysis infallible — check anything important against the recording.
📖 History
Which model to pick, and what it costs you
Recognition runs in your browser, so the model is downloaded once and cached. Bigger models hear accents and noise better but take longer on the same machine.
| Model | One-time download | Time per minute of audio | When to pick it |
|---|---|---|---|
| Fast | ~75 MB | about 1 min | Quick draft, clear speech, one speaker |
| Accurate | ~150 MB | about 3 min | Everyday recordings — the usual compromise |
| High Accuracy (small) | ~250 MB | about 6 min | Accents, background noise, several speakers |
| Premium (Large v3 Turbo) | ~600 MB | depends on the GPU | Hardest audio; needs a GPU or a modern CPU |
A video file is stripped to its audio track before anything is processed: an hour of phone video is 2–3 GB, of which the sound is about 30 MB.