Video OCR: Extract Text from a Video
Read the text shown in a video: slides, code and captions, with a timestamp for each scene.
Reads the frames: scenes and the text on screen. One minute per minute of video.
Free to try, no sign-up. Watch is one minute per minute of video. Both is two.
Have a file instead? Reading the screen of a local file runs in the command-line tool or the Mac app.
How it works
- 1
Paste a link
A public video link. For a file, use the command-line tool or the Mac app.
- 2
Scribiz reads the frames
Frames are sampled and grouped into scenes, and the text in each is read.
- 3
Copy the text
The text of each scene with its timestamp. Export Context or JSON.
On-screen text, scene by scene
- 0:00
A title slide on a dark background.
On screenShipping faster with fewer meetingsTeam offsite, day 2
- 0:42
A slide with three numbered points.
On screen1. Write it down2. Decide in the doc3. Meet only to unblock
- 1:35
A terminal window with one command.
On screengit log --since="1 week ago" --oneline
Text that is in the picture, not in the audio
A transcript holds what was said. A lot of video says the important part in writing: the slide behind the speaker, the command in a terminal, the price on an ad, the caption drawn over a Reel. Video OCR reads that text off the frames.
Pick Watch to read the screen only, or Watch and listen to get the speech beside it. The result opens on the On screen tab: one row per scene, with its time, a line on what the picture shows and the text that was visible.
| On screen | What comes back |
|---|---|
| Slides | Titles and bullet text, one scene per slide. |
| Code and terminals | The visible lines. Long files that scroll come back in parts. |
| Burned-in captions | The words drawn on the picture, grouped by scene. |
| Lower thirds and labels | Names, titles and prices as they appear. |
Video OCR compared with OCR on a picture
Classic OCR takes one still image. A video is thousands of them, mostly repeats. Scribiz looks at a frame every second or two, groups frames that show the same thing and writes the text once per scene, so a slide that stays up for a minute is one row and not sixty.
Text that is on screen for less than a second can fall between two frames. Very small type, handwriting and text at an angle are the usual misses. The text comes from a model that describes what it sees, so copy code and figures with care.
Export the text from a video
Copy the rows from the On screen tab, or export. The Context and JSON files carry the on-screen text with its timestamps, which is the form to paste into an AI chat or feed to a script. The Context object describes the fields.
Watch is one minute per minute of video and Watch and listen is two. If you want a description of each scene more than its text, the video analyzer is the same run read the other way. For the spoken words alone, use video to text.
Limits
- Public links on the web. Reading the screen of a file runs in the command-line tool or the Mac app.
- Small or fast text can be misread. Check names, numbers and code.
- Without an account: videos up to 15 minutes and 10 Watch minutes a day.
Frequently asked questions
Transcripts in agents and code
Use it in Claude, Cursor or Codex with the MCP server, or from your own code with the API.
