Skip to content
Try it free

Start

Overview

What a Context is, the four layers it holds, and how Auto, Listen, Watch and Both decide what runs.

View as Markdown
On this page

Scribiz turns a video into a Context: what was said, what was shown, and what it adds up to. You give it a link or a file. You get one object that you can read, search, export or hand to an agent.

The same Context comes out of the web tool, the CLI, the MCP server and the API.

What a Context holds

A Context has four layers.

LayerWhat it isWhere it comes from
TranscriptThe spoken words with timestamps, and speakers when there is more than one voiceThe platform's captions, or speech to text on the audio
On screenScenes, and the text visible in each oneA vision model reading a low-resolution copy of the video
SummaryA short summary, key points and key momentsWritten from the transcript and the on-screen notes
ChaptersTitled sections that start at real points in the transcriptWritten from the transcript, then snapped to transcript segments

Every layer says how it was made. A transcript carries its origin (captions or speech to text) and its timing quality. A Context lists warnings when something is approximate or incomplete. Read The Context object for every field.

Times are never taken on trust. A model decides where a chapter or a citation belongs, and Scribiz then moves it to the nearest transcript segment, so the timestamp lands where the words are.

Modes

A mode decides which layers run. You can set it on the web tool, with --mode in the CLI, or with mode in the API.

ModeFlag valueWhat runsMinutes per minute of video
AutoautoThe cheapest path that works, see below0 to 2, depending on the path
Listenaudio (alias listen)Speech to text on the audio1
Watchvisual (alias watch)Scenes and on-screen text1
Bothfull (alias both)Listen and Watch2
CaptionscaptionsPlatform captions only0.1

Captions use a tenth of a minute for each minute of video. Without an account they are limited to 30 lookups a day.

The Summary and Chapters layers are written last in every mode, from whatever the other layers found. Without a summary, as with --detail transcript in the CLI, a run stops at the transcript.

Minutes and billing explains how those numbers add up.

What Auto does

Auto looks at the video first and then picks, in this order:

  1. A stored result for the same video, when there is one. It costs nothing.
  2. The platform's own captions, when the video has a good set: a manual track in the video's language, or the automatic track in the original language when it passes a quality check. Machine translations of captions are skipped.
  3. Listen, when there are no usable captions or you asked for speakers.
  4. Watch as well, when the video has no audio, when there is almost no speech (music, a silent screen recording), or when it is a short clip, up to three minutes, from Instagram, TikTok or YouTube Shorts.

A long talk with normal speech never gets Watch in Auto. Ask for Watch or Both when you want it.

If a public YouTube video has no captions and its audio cannot be fetched, Scribiz can have the model read the video link directly. The transcript then has approximate timing and no speaker labels, and a warning (DOWNLOAD_FALLBACK_URL_DIRECT) says so.

Where to run it

SurfaceUse it forStart here
WebOne-off links and files, nothing to installQuickstart
CLIScripts, batch jobs, editors, local files, your own IP and browser cookies, with a free Scribiz account or your own Gemini keyInstall the CLI
Mac app, not published yetLocal files, folders and recordings, with your own key or a Scribiz accountMac app
MCP serverLetting Claude, Cursor, Codex or any MCP client work from the videoMCP overview
APIYour own productAPI overview

Instagram, TikTok and other sites that block servers work best from the CLI, because it fetches from your own connection. It is on npm (npm install -g scribiz) and has been tested on macOS only. The Mac app, which will do the same, is not published yet. Sources and limits has the full matrix.

Next

Loading the index