Transcribe a YouTube Video to Text
Real speech-to-text from the audio. Auto-captions not required.
Transcribes the audio. One minute per minute of video.
Listen ignores the video captions and transcribes the speech. One minute per minute of video.
How it works
Paste the link
Any public YouTube video, with or without captions.
Scribiz listens
Listen mode skips the captions and transcribes the sound. If you want captions first, use the YouTube transcript page.
Check the names
Skim names, numbers and terms, then copy or export.
Jargon from the audio
- 0:00
Before we touch the benchmark, a word about what we are measuring: p99 latency, not the average.
- 0:07
The average hides the slow requests, and the slow requests are what users remember.
- 0:14
We ran nginx in front of two Kubernetes pods and replayed an hour of real traffic.
- 0:22
Notice the spike at minute forty. That is the garbage collector, not the network.
Illustrative sample from a technical talk. Terms and numbers come from the sound, not from captions.
When auto-captions get it wrong
Auto-captions stumble on accents, jargon, crosstalk and music. A lecture with a strong accent, a podcast where two people talk over each other and a talk full of product names are the usual cases. Listen mode ignores the captions and transcribes the sound.
It is a fresh attempt, not a perfect one. Names and numbers still need a read. It also covers videos where the uploader turned captions off. The guide to a transcript with no captions shows how to check that first.
Captions or listening
This page runs Listen, so the video captions are ignored. The cheaper run is the YouTube transcript page, which uses good captions and listens only when it has to.
Listening is worth it when you will quote people, publish the text or hand it to an AI, because a wrong term spreads. For a quick skim, captions are enough. Listen and watch adds the picture, for talks where the slides carry the terms.
| Run | Where the text comes from | Minutes per minute of video |
|---|---|---|
| Captions | The track already on the video | A tenth |
| Listen | The audio, transcribed fresh | One |
| Listen and watch | The audio and the frames | Two |
Timing and speakers
From a YouTube link, timestamps are per sentence and there are no speaker labels. YouTube blocks our server from downloading the audio, so Google's model reads the public link instead. That path gives approximate timing, and the result page says so.
For word timing and Speaker 1, Speaker 2, drop the original recording on the audio to text page. The order Scribiz tries for a YouTube link is in sources and limits.
Limits
Public videos only. Timestamps from a link can drift by a couple of seconds. Without an account: videos up to 15 minutes, 10 Listen or Watch minutes a day and 30 caption lookups a day. Listening uses one minute per minute of video.
See pricing · How minutes are counted · Paid plans are not on sale yet.