How to transcribe audio to text
- Drop the recording on the box above, or press Choose file. Leave the language on Auto-detect unless the recording starts with music or is very short.
- Wait while it reads the file and transcribes the speech. Most recordings are done in under a minute.
- Play the audio and read along. The line being spoken is shown above the timeline; click any line in the list to jump to it and fix a word.
- The Export tab is already open. Press TXT for a plain transcript, or SRT or VTT if you need timed subtitles.
That's the whole job, and the editor keeps going after the first download. Fix more lines and download again as often as you like, without spending more credits.
Free audio transcription, with limits stated up front
Every visitor gets 20 credits a day, and transcription costs 1 credit per started minute of audio, so a 12-minute interview uses 12. Files can be up to 30 minutes long. Credits refill daily, and everything after the transcription (editing, downloading TXT, SRT or VTT, starting over with the same transcript) is free.
The speech model is Whisper large-v3 turbo, running on Groq's hardware, which is why most files come back in well under a minute. Your browser opens the file, makes a small compressed mono copy of the audio, and sends only that copy. We don't keep it or the text. Our article on how accurate Whisper transcription is covers where it does well and where it slips.
MP3 to text, WAV, M4A and other formats
MP3 to text is the most common job, but the tool takes most audio you're likely to have: M4A from an iPhone voice memo or a Zoom audio-only recording, WAV from a recorder or a DAW export, AAC, OGG, Opus and FLAC. Big WAV files are fine because the file never uploads; only the compressed speech copy does.
The file is decoded by your browser. Current Chrome and Edge open everything above. If a file won't open in another browser, re-save it as MP3 or WAV, or switch to Chrome. If what you have is a video file, drop it in all the same, or use the video to text page, which explains that side in more detail.
Podcast transcription
A transcript makes an episode searchable and gives you show notes and quotes without listening twice. Export TXT for the notes, or SRT if you publish the episode as a video on YouTube and want proper captions there. Two things help. Transcribe the final mix rather than the raw tracks, so the timing matches what listeners hear. And for episodes longer than 30 minutes, split the file at a natural break; each part's timestamps start from zero.
Interview transcription
For research interviews, journalism and user calls, the useful thing is finding a quote fast. Play the recording, click the line where the quote starts, and the audio jumps there. The tool doesn't label speakers, so add names as you read. Changing a line's text doesn't change its timing. If you record with a phone, put it close to the person you need to hear clearly; distance and room echo cost far more accuracy than the phone itself.
Lecture and voice memo transcription
Lectures run long, so check the length first. Anything up to 30 minutes goes in one go. Voice memos are the easy case: usually one speaker, close to the mic, and short enough to cost a few credits. Pick the language yourself if the memo mixes languages or opens with a few seconds of silence, because auto-detect listens to the start of the file.
What you get back
| Format | What it is | Use it for |
|---|---|---|
| TXT | The words, one line per caption, no timestamps | Notes, articles, quotes, searching |
| SRT | Numbered, timed subtitle blocks | YouTube, video editors, social platforms |
| VTT | The web's caption format | Your own site's video or audio player |
| SBV | YouTube's caption format | Older YouTube workflows |
Need to change format later? The subtitle converter turns any of these timed files into the others.