Free ยท MP3, M4A, WAV ยท No sign-up

Audio to text, free and online

Drop in a podcast episode, interview, lecture or voice memo and get a timed transcript back in 50+ languages. Read it, fix it while you listen, then copy it or download TXT, SRT or VTT.

Drop a video or audio file

MP4, MOV or WebM video, or MP3, M4A or WAV audio, up to 30 minutes

Already have subtitles? Drop the SRT or VTT in too, or on its own to edit it.

Your video stays on your device. Only a compressed copy of the audio is sent for transcription.

Captions cost 1 credit per minute of video or audio.Importing subtitles is free.

  • Free, no watermark
  • No sign-up
  • MP4, SRT, VTT or TXT
  • 50+ languages

How to transcribe audio to text

  1. Drop the recording on the box above, or press Choose file. Leave the language on Auto-detect unless the recording starts with music or is very short.
  2. Wait while it reads the file and transcribes the speech. Most recordings are done in under a minute.
  3. Play the audio and read along. The line being spoken is shown above the timeline; click any line in the list to jump to it and fix a word.
  4. The Export tab is already open. Press TXT for a plain transcript, or SRT or VTT if you need timed subtitles.

That's the whole job, and the editor keeps going after the first download. Fix more lines and download again as often as you like, without spending more credits.

Free audio transcription, with limits stated up front

Every visitor gets 20 credits a day, and transcription costs 1 credit per started minute of audio, so a 12-minute interview uses 12. Files can be up to 30 minutes long. Credits refill daily, and everything after the transcription (editing, downloading TXT, SRT or VTT, starting over with the same transcript) is free.

The speech model is Whisper large-v3 turbo, running on Groq's hardware, which is why most files come back in well under a minute. Your browser opens the file, makes a small compressed mono copy of the audio, and sends only that copy. We don't keep it or the text. Our article on how accurate Whisper transcription is covers where it does well and where it slips.

MP3 to text, WAV, M4A and other formats

MP3 to text is the most common job, but the tool takes most audio you're likely to have: M4A from an iPhone voice memo or a Zoom audio-only recording, WAV from a recorder or a DAW export, AAC, OGG, Opus and FLAC. Big WAV files are fine because the file never uploads; only the compressed speech copy does.

The file is decoded by your browser. Current Chrome and Edge open everything above. If a file won't open in another browser, re-save it as MP3 or WAV, or switch to Chrome. If what you have is a video file, drop it in all the same, or use the video to text page, which explains that side in more detail.

Podcast transcription

A transcript makes an episode searchable and gives you show notes and quotes without listening twice. Export TXT for the notes, or SRT if you publish the episode as a video on YouTube and want proper captions there. Two things help. Transcribe the final mix rather than the raw tracks, so the timing matches what listeners hear. And for episodes longer than 30 minutes, split the file at a natural break; each part's timestamps start from zero.

Interview transcription

For research interviews, journalism and user calls, the useful thing is finding a quote fast. Play the recording, click the line where the quote starts, and the audio jumps there. The tool doesn't label speakers, so add names as you read. Changing a line's text doesn't change its timing. If you record with a phone, put it close to the person you need to hear clearly; distance and room echo cost far more accuracy than the phone itself.

Lecture and voice memo transcription

Lectures run long, so check the length first. Anything up to 30 minutes goes in one go. Voice memos are the easy case: usually one speaker, close to the mic, and short enough to cost a few credits. Pick the language yourself if the memo mixes languages or opens with a few seconds of silence, because auto-detect listens to the start of the file.

What you get back

FormatWhat it isUse it for
TXTThe words, one line per caption, no timestampsNotes, articles, quotes, searching
SRTNumbered, timed subtitle blocksYouTube, video editors, social platforms
VTTThe web's caption formatYour own site's video or audio player
SBVYouTube's caption formatOlder YouTube workflows

Need to change format later? The subtitle converter turns any of these timed files into the others.

FAQ

Questions, answered

01

Which audio formats can I transcribe?

MP3, M4A, WAV, AAC, OGG, Opus and FLAC, plus voice memos saved as M4A. Your browser decodes the file, so if one refuses to open, re-save it as MP3 or WAV, or try Chrome or Edge.
02

Is it really free?

Yes. Transcription uses credits that refill every day: 1 credit per started minute, and every visitor gets 20 a day. Editing the transcript and downloading it again cost nothing. There is no watermark and no account.
03

How long can the recording be?

Up to 30 minutes per file. For a longer podcast or lecture, split it into parts in any audio editor and transcribe them one at a time, or come back the next day when your credits refill.
04

Does it label who is speaking?

No. The transcript is the words in order, timed line by line, without speaker names. For a two-person interview it's usually quick to add names while you read through, because every line can be edited.
05

What happens to my audio?

Your browser makes a small compressed copy of the audio and sends only that to the speech model (Whisper large-v3 turbo, running on Groq). We do not store the audio or the transcript that comes back.
06

How accurate is the transcript?

Clear speech with one person close to the microphone comes back nearly clean. Music, crosstalk, heavy accents and names cause most of the errors, which is why the transcript opens in an editor where you can fix a line while listening to it.

Still have a question? Contact us

Keep going

Related tools and guides

Tool

Video to text

Convert video to text free. Whisper transcribes MP4, MOV and WebM in 50+ languages. Copy the transcript or download TXT, SRT or VTT. No sign-up.

Tool

SRT generator

Free SRT generator. Turn an MP4 or MOV into a timed SRT or VTT subtitle file in seconds, fix any line, and download it. No sign-up, no watermark.

Blog

Whisper AI transcription

How Whisper AI transcription works, where it gets captions wrong, what word-level timestamps are, and how to check and fix auto captions quickly.

Try it now

Transcribe your recording now

Free, no sign-up. Only a compressed copy of the audio is sent for transcription.