Skip to content
Get transcript

Video OCR: Extract Text from a Video

Read the text shown in a video: slides, code and captions, with a timestamp for each scene.

Reads the frames: scenes and the text on screen. One minute per minute of video.

Free to try, no sign-up. Watch is one minute per minute of video. Both is two.

Have a file instead? Reading the screen of a local file runs in the command-line tool or the Mac app.

Open a recorded example

How it works

  1. 1

    Paste a link

    A public video link. For a file, use the command-line tool or the Mac app.

  2. 2

    Scribiz reads the frames

    Frames are sampled and grouped into scenes, and the text in each is read.

  3. 3

    Copy the text

    The text of each scene with its timestamp. Export Context or JSON.

On-screen text, scene by scene

Sample
  1. 0:00

    A title slide on a dark background.

    On screenShipping faster with fewer meetingsTeam offsite, day 2

  2. 0:42

    A slide with three numbered points.

    On screen1. Write it down2. Decide in the doc3. Meet only to unblock

  3. 1:35

    A terminal window with one command.

    On screengit log --since="1 week ago" --oneline

Illustrative sample from a slide talk. The text is read by a model, so treat it as a draft, not a record.

Text that is in the picture, not in the audio

A transcript holds what was said. A lot of video says the important part in writing: the slide behind the speaker, the command in a terminal, the price on an ad, the caption drawn over a Reel. Video OCR reads that text off the frames.

Pick Watch to read the screen only, or Watch and listen to get the speech beside it. The result opens on the On screen tab: one row per scene, with its time, a line on what the picture shows and the text that was visible.

On screenWhat comes back
SlidesTitles and bullet text, one scene per slide.
Code and terminalsThe visible lines. Long files that scroll come back in parts.
Burned-in captionsThe words drawn on the picture, grouped by scene.
Lower thirds and labelsNames, titles and prices as they appear.

Video OCR compared with OCR on a picture

Classic OCR takes one still image. A video is thousands of them, mostly repeats. Scribiz looks at a frame every second or two, groups frames that show the same thing and writes the text once per scene, so a slide that stays up for a minute is one row and not sixty.

Text that is on screen for less than a second can fall between two frames. Very small type, handwriting and text at an angle are the usual misses. The text comes from a model that describes what it sees, so copy code and figures with care.

Export the text from a video

Copy the rows from the On screen tab, or export. The Context and JSON files carry the on-screen text with its timestamps, which is the form to paste into an AI chat or feed to a script. The Context object describes the fields.

Watch is one minute per minute of video and Watch and listen is two. If you want a description of each scene more than its text, the video analyzer is the same run read the other way. For the spoken words alone, use video to text.

Limits

  • Public links on the web. Reading the screen of a file runs in the command-line tool or the Mac app.
  • Small or fast text can be misread. Check names, numbers and code.
  • Without an account: videos up to 15 minutes and 10 Watch minutes a day.

See pricing · How minutes are counted

Frequently asked questions

Reading the text that is visible in the frames of a video: slides, code, captions drawn on the picture, labels and prices. It is separate from the transcript, which is what was said.

Printed text of a readable size that stays up for a second or more is read well. Small type, handwriting, text at an angle and text that flashes by are the usual misses. A model reads the frames, so check code and numbers.

Not in the browser, where a file is read for its audio only. Paste a public link here, or use the command-line tool or the Mac app for a file.

Yes. Pick Watch. A silent screen recording or a slideshow with music gives the on-screen text and nothing in the transcript.

Each scene has the time it starts. A slide that stays on screen is one scene, so you get its text once and not once per frame.

15 minutes without an account, 2 hours with a free account, 6 hours on Pro. Watch counts as 1 minute per minute of video and Watch and listen as 2.

Transcripts in agents and code

Use it in Claude, Cursor or Codex with the MCP server, or from your own code with the API.