Skip to content

Audio to text converter

Open an MP3, WAV, M4A or voice note and read it as text a moment later. The file is converted on your device, not on a server, and there is nothing to install.

  • Free
  • No sign-up
  • No upload
Saves asSRTVTTTXTDOCX

Drop an audio or video file here

MP3, WAV, M4A, OGG, Opus, FLAC, WebM, MP4, MOV

Formats it reads

The page hands your file to the browser’s own audio decoder, the same one that plays sound on any website. Whatever it can decode, the speech model can read.

MP3Podcasts, voice recorders, downloads
Opens everywhere. 64 kbps mono is already enough for speech.
M4A / AACiPhone Voice Memos, Zoom and Teams recordings
Opens in current browsers. Long meeting recordings are best split into parts.
WAVStudio recorders, audio editors
Uncompressed, so files are large: about 10 MB per minute in CD quality.
OGG / OpusWhatsApp and Telegram voice notes
Opens directly; no conversion to MP3 needed.
FLACArchival copies
Lossless. Opens in Chrome, Edge, Firefox and current Safari.
WebMRecordings made in a browser
Usually Opus inside. Opens in Chrome, Edge and Firefox.

What actually improves a transcript

The model listens to a 16 kHz mono version of your file, close to telephone quality. That is why bitrate and sample rate barely matter, and why the recording itself matters a great deal.

  • Distance. A phone on the table half a metre from the speaker beats an expensive microphone across the room.
  • One voice at a time. Overlapping speech is where every speech model loses words.
  • Background. Music, traffic and fans turn into wrong words, or into text that was never said. Stretches of pure silence are skipped on purpose, because Whisper tends to invent sentences for them.
  • Tell it the language. Automatic detection listens to the first thirty seconds. If a recording starts with music or with another language, choose the language yourself.

Names, brands and technical terms are the usual leftovers. Fix them once in the editor before exporting, or tidy dictated text on the dictation page.

Once you have the text

Copy text gives you plain paragraphs for an email or a document. TXT and DOCX keep the start time of each line, which is what you want when quoting an interview or citing a lecture. SRT and VTT are subtitle files; the SRT generator page explains how to time them well.

Questions and answers

Which audio formats can I open?

Whatever your browser can decode: MP3, WAV, M4A (AAC), OGG and Opus work in current Chrome, Edge, Firefox and Safari; FLAC and WebM work in most. If a file cannot be read, the page says so and nothing else happens.

Do I need to convert my file to MP3 first?

No. The browser decodes the file to plain 16 kHz mono sound, which is what the speech model reads. Converting beforehand only costs time and quality.

Does a higher bitrate give a better transcript?

Only up to a point. The model listens to 16 kHz mono, so a 320 kbps file is no better than a 96 kbps one. What matters is the recording: distance from the microphone, background noise and people talking at once.

Can I get the text without timestamps?

“Copy text” copies the transcript as plain paragraphs. The TXT and DOCX files include the start time of each line, which helps when you quote from an interview.

Is there a limit on how many files I can transcribe?

No. Nothing is counted because nothing is sent to us: the work happens on your device.