Audio to text converter
Open an MP3, WAV, M4A or voice note and read it as text a moment later. The file is converted on your device, not on a server, and there is nothing to install.
- Free
- No sign-up
- No upload
Drop an audio or video file here
MP3, WAV, M4A, OGG, Opus, FLAC, WebM, MP4, MOV
Formats it reads
The page hands your file to the browser’s own audio decoder, the same one that plays sound on any website. Whatever it can decode, the speech model can read.
- MP3Podcasts, voice recorders, downloads
- Opens everywhere. 64 kbps mono is already enough for speech.
- M4A / AACiPhone Voice Memos, Zoom and Teams recordings
- Opens in current browsers. Long meeting recordings are best split into parts.
- WAVStudio recorders, audio editors
- Uncompressed, so files are large: about 10 MB per minute in CD quality.
- OGG / OpusWhatsApp and Telegram voice notes
- Opens directly; no conversion to MP3 needed.
- FLACArchival copies
- Lossless. Opens in Chrome, Edge, Firefox and current Safari.
- WebMRecordings made in a browser
- Usually Opus inside. Opens in Chrome, Edge and Firefox.
What actually improves a transcript
The model listens to a 16 kHz mono version of your file, close to telephone quality. That is why bitrate and sample rate barely matter, and why the recording itself matters a great deal.
- Distance. A phone on the table half a metre from the speaker beats an expensive microphone across the room.
- One voice at a time. Overlapping speech is where every speech model loses words.
- Background. Music, traffic and fans turn into wrong words, or into text that was never said. Stretches of pure silence are skipped on purpose, because Whisper tends to invent sentences for them.
- Tell it the language. Automatic detection listens to the first thirty seconds. If a recording starts with music or with another language, choose the language yourself.
Names, brands and technical terms are the usual leftovers. Fix them once in the editor before exporting, or tidy dictated text on the dictation page.
Once you have the text
Copy text gives you plain paragraphs for an email or a document. TXT and DOCX keep the start time of each line, which is what you want when quoting an interview or citing a lecture. SRT and VTT are subtitle files; the SRT generator page explains how to time them well.
Questions and answers
Which audio formats can I open?
Whatever your browser can decode: MP3, WAV, M4A (AAC), OGG and Opus work in current Chrome, Edge, Firefox and Safari; FLAC and WebM work in most. If a file cannot be read, the page says so and nothing else happens.
Do I need to convert my file to MP3 first?
No. The browser decodes the file to plain 16 kHz mono sound, which is what the speech model reads. Converting beforehand only costs time and quality.
Does a higher bitrate give a better transcript?
Only up to a point. The model listens to 16 kHz mono, so a 320 kbps file is no better than a 96 kbps one. What matters is the recording: distance from the microphone, background noise and people talking at once.
Can I get the text without timestamps?
“Copy text” copies the transcript as plain paragraphs. The TXT and DOCX files include the start time of each line, which helps when you quote from an interview.
Is there a limit on how many files I can transcribe?
No. Nothing is counted because nothing is sent to us: the work happens on your device.
More tools on this site
- Transcribe audio to textOpen a recording, get the text, correct it while you listen.Open
- Video to textTranscribe the audio track and preview subtitles over the picture.Open
- SRT generatorTimed subtitles from a recording, exported as SRT or VTT.Open
- Subtitle editorFix an existing .srt or .vtt file, with or without the video.Open
- DictationSpeak and get text, with a local option that keeps the audio here.Open
- Transcrever áudio em textoA mesma ferramenta, em português.Abrir
- Transcrever áudio do WhatsAppArquivos .opus direto, sem converter.Abrir