Skip to content
Try it free

Audio to Text Converter

Drop an audio file. Get a transcript with speakers and timestamps.

or

Free to try, no sign-up. You can also paste a direct link to an audio file or a podcast episode.

How it works

  1. Choose a file

    MP3, M4A, WAV, OGG, FLAC or AAC. It stays in your browser until you press the button.

  2. Check the minutes

    The page shows what the run will use first.

  3. Read and export

    Speaker-labeled paragraphs with timestamps, or SRT, VTT, TXT, Markdown and JSON.

What to expect

Sample
  1. 0:02

    Speaker 1

    Thanks for joining. Let us start with how the bike lane project began.

  2. 0:08

    Speaker 2

    It began with a petition. Three hundred signatures, mostly from parents on one school route.

  3. 0:16

    Speaker 1

    And the council said yes right away?

  4. 0:19

    Speaker 2

    No. They asked for a traffic count first, and that took until spring.

Illustrative, not a real recording. Voices come back as Speaker 1 and Speaker 2.

Speakers and timestamps

Each paragraph gets a timestamp and a label. Two voices work best: an interview comes back clean, and a meeting with several voices gets useful text with some speaker swaps.

You can rename Speaker 1 on the result page, and the new name carries into the exports.

RecordingWhat to expect
An interview with two peopleClean turns, labelled Speaker 1 and Speaker 2.
A lecture or talk, one voiceParagraphs with timestamps. Check names and terms.
A meeting with four or more voicesUseful text, with some speaker swaps.
A phone call or voicemailReadable when the line is clear. Hold music and crosstalk cost quality.

Noise, accents and names

A phone held near the mouth beats an expensive microphone across the room. Music under speech and crosstalk cause most errors.

Names and brand names are the common slip. Skim for them. Trim long silences and hold music if you can: the length of the file, not its size, sets the minutes a run uses.

Formats and long recordings

The browser sends an audio file without converting it, and our server then checks that it can read it. MP3 and M4A have their own pages, MP3 to text and M4A to text. A video file goes to video to text, which reads the audio in your browser first.

Files go up to 100 MB in the browser, and your plan sets the length. For subtitles, run the file and export SRT or VTT, or see the SRT generator. The API quickstart shows how to send a file from code.

Limits

100 MB per file in the browser. Without an account: files up to 15 minutes and 10 minutes a day. A free account takes recordings up to 2 hours and 30 minutes a month. Pro takes 6 hours and 600 minutes a month, and is not on sale yet. Speaker labels are experimental for three or more voices.

See pricing · How minutes are counted · Paid plans are not on sale yet.

Questions

MP3, M4A, WAV, OGG, FLAC, AAC and Opus, up to 100 MB in the browser. An audio file is uploaded as it is, and our server checks that it can read it.

Clean recordings come back clean. Noise, crosstalk and music under the voice cost quality, and uncommon names are often misspelled.

Yes, as Speaker 1, Speaker 2 and so on. Two voices work best, and three or more can get mixed up.

The browser takes 100 MB. Length limits: 15 minutes without an account, 2 hours with a free account, 6 hours on Pro. Minutes and billing has the details.

Yes, when the link points straight at the audio file, such as an episode URL ending in .mp3. Show pages and copy-protected services do not work, so drop the file instead.

The audio is deleted when the job succeeds, or within an hour if it fails. The transcript stays on a private link for 24 hours, or 30 days with an account.