Skip to content
Try it free

Transcribe a YouTube Video to Text

Real speech-to-text from the audio. Auto-captions not required.

Transcribes the audio. One minute per minute of video.

Listen ignores the video captions and transcribes the speech. One minute per minute of video.

Open a recorded example

How it works

  1. Paste the link

    Any public YouTube video, with or without captions.

  2. Scribiz listens

    Listen mode skips the captions and transcribes the sound. If you want captions first, use the YouTube transcript page.

  3. Check the names

    Skim names, numbers and terms, then copy or export.

Jargon from the audio

Sample
  1. 0:00

    Before we touch the benchmark, a word about what we are measuring: p99 latency, not the average.

  2. 0:07

    The average hides the slow requests, and the slow requests are what users remember.

  3. 0:14

    We ran nginx in front of two Kubernetes pods and replayed an hour of real traffic.

  4. 0:22

    Notice the spike at minute forty. That is the garbage collector, not the network.

Illustrative sample from a technical talk. Terms and numbers come from the sound, not from captions.

When auto-captions get it wrong

Auto-captions stumble on accents, jargon, crosstalk and music. A lecture with a strong accent, a podcast where two people talk over each other and a talk full of product names are the usual cases. Listen mode ignores the captions and transcribes the sound.

It is a fresh attempt, not a perfect one. Names and numbers still need a read. It also covers videos where the uploader turned captions off. The guide to a transcript with no captions shows how to check that first.

Captions or listening

This page runs Listen, so the video captions are ignored. The cheaper run is the YouTube transcript page, which uses good captions and listens only when it has to.

Listening is worth it when you will quote people, publish the text or hand it to an AI, because a wrong term spreads. For a quick skim, captions are enough. Listen and watch adds the picture, for talks where the slides carry the terms.

RunWhere the text comes fromMinutes per minute of video
CaptionsThe track already on the videoA tenth
ListenThe audio, transcribed freshOne
Listen and watchThe audio and the framesTwo

Timing and speakers

From a YouTube link, timestamps are per sentence and there are no speaker labels. YouTube blocks our server from downloading the audio, so Google's model reads the public link instead. That path gives approximate timing, and the result page says so.

For word timing and Speaker 1, Speaker 2, drop the original recording on the audio to text page. The order Scribiz tries for a YouTube link is in sources and limits.

Limits

Public videos only. Timestamps from a link can drift by a couple of seconds. Without an account: videos up to 15 minutes, 10 Listen or Watch minutes a day and 30 caption lookups a day. Listening uses one minute per minute of video.

See pricing · How minutes are counted · Paid plans are not on sale yet.

Questions

Auto-captions are speech recognition YouTube ran once, with no review. Transcribing means Scribiz listens to the audio and writes a fresh transcript.

Clear speech comes back clean. Crosstalk, music, heavy accents and specialist terms cause errors, so read it before you quote it.

Not from a YouTube link on the web. Drop the audio on the audio to text page to get Speaker 1, Speaker 2 and so on.

Scribiz writes the language that is spoken, and widely spoken languages work best. It does not translate yet.

No. You paste a link. Listen skips the captions: YouTube blocks our server from downloading the audio, so Google's model reads the public link and writes the text. Timing is approximate, to the sentence, and speakers are not labelled. For captions first, use the YouTube transcript page.

Listening uses one minute per minute of video, and without an account you get 10 a day for videos up to 15 minutes. Captions cost a tenth of a minute per minute of video, so the YouTube transcript page is the cheaper run when they are good enough.