API
The Context object
Every field of the Context: metadata, transcript, on-screen notes, summary, chapters, coverage, warnings and usage.
On this page
The Context is what every surface returns. The API puts it in result. The CLI prints it with --format json. The MCP server returns parts of it.
An example
An abridged Context for a 19 second public YouTube video, as the API returns it in result. Long fields are shortened and the list of segments is cut to two.
{ "schema": 1, "id": "youtube:jNQXAC9IVRw", "source": { "kind": "url", "provider": "youtube", "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw", "canonicalUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw", "videoId": "jNQXAC9IVRw", "extractor": "Youtube" }, "meta": { "title": "Me at the zoo", "durationSeconds": 19, "hasAudio": true, "hasVideo": true, "channel": { "id": "UC4QobU6STFB0P71PMvOGN5A", "name": "jawed", "url": "https://www.youtube.com/channel/UC4QobU6STFB0P71PMvOGN5A" }, "uploadDate": "2005-04-24", "chapters": [ { "start": 0, "end": 5, "title": "Intro" }, { "start": 5, "end": 17, "title": "The cool thing" }, { "start": 17, "end": 19, "title": "End" } ], "width": 320, "height": 240, "fps": 15 }, "transcript": { "origin": "captions-manual", "language": "en", "timing": "caption", "diarized": false, "proofread": "none", "speakers": [], "segments": [ { "start": 1.2, "end": 3.36, "text": "All right, so here we are, in front of the elephants" }, { "start": 5.318, "end": 7.974, "text": "the cool thing about these guys is that they have really..." } ], "text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... ..." }, "visual": null, "synthesis": { "model": "gemini-3.8-flash", "language": "en", "title": "Observing Elephants at the Zoo", "summary": { "tldr": "The speaker stands in front of elephants and points out that they have very long trunks.", "bullets": ["The speaker is positioned directly in front of the elephants.", "..."], "long": "..." }, "chapters": [], "keyMoments": [ { "time": 1.2, "label": "Arrival in front of the elephants", "kind": "other", "segmentIndex": 0 }, { "time": 5.318, "label": "Remarking on the elephants' really long trunks", "kind": "claim", "segmentIndex": 1 } ], "entities": [{ "name": "elephants", "kind": "term", "mentions": 1 }], "topics": ["elephants", "elephant trunks", "zoo animals"] }, "coverage": { "speechSeconds": null, "transcribedSeconds": 14.521, "ratio": null, "repairedWindows": 0, "unrepairedWindows": 0, "gaps": [] }, "tiers": [ { "tier": "captions", "status": "ok", "reason": "manual en captions (the video language is unknown)", "ms": 5472, "usd": 0 }, { "tier": "visual", "status": "skipped", "reason": "not needed: normal speech density", "ms": 0, "usd": 0 }, { "tier": "synthesis", "status": "ok", "ms": 2717, "retries": 0, "usd": 0.0017175 } ], "warnings": [], "usage": { "audioSeconds": 0, "videoSeconds": 0, "tokens": { "in": 570, "out": 344, "thought": 0, "audioIn": 0, "videoIn": 0 }, "usdEstimate": 0.0017175 }, "produced": { "at": "2026-10-04T05:57:33.671Z", "engine": "0.1.0", "models": { "synthesis": "gemini-3.8-flash" }, "promptVersion": 1, "mode": "auto" }}A Context for a recording that was listened to has speakers and, in the CLI, word times:
{ "origin": "asr", "model": "gemini-3.5-transcribe", "language": "en", "timing": "word", "diarized": true, "speakers": [ { "id": "S1", "wordCount": 15, "seconds": 5 }, { "id": "S2", "wordCount": 16, "seconds": 5 } ], "segments": [ { "start": 0.3, "end": 2.4, "speaker": "S1", "text": "Welcome to the Scribus fixture test.", "wordRange": [0, 6] } ], "words": [ { "start": 0.3, "end": 0.7, "speaker": "S1", "text": "Welcome" } ]}A Context for a video that was watched has a visual object. This one is from an 8 second clip with no sound, where transcript is null and the warning NO_SPEECH_DETECTED is set:
{ "model": "gemini-3.5-flash-lite", "source": "proxy-upload", "fps": 2, "value": "high", "overview": "The video displays three consecutive illustrated scenes featuring different objects and text titles, representing a red apple, a green forest, and a blue ocean.", "scenes": [ { "start": 0, "end": 3, "description": "A cream-colored background shows a red circle resembling an apple in the center, with text at the bottom.", "onScreenText": ["SCENE ONE - RED APPLE"], "timingApprox": true } ], "onScreenText": [{ "text": "SCENE ONE - RED APPLE", "first": 0, "last": 3 }]}Reading it safely
- All times are seconds, as numbers, rounded to milliseconds.
- The Context has a
schemanumber. Fields are only ever added within a schema. Ignore fields you do not know. transcriptisnullonly when the video has no speech.visualisnullwhen Watch did not run.synthesisisnullwhen the summary step did not run.- Check
transcript.timingbefore you do anything that needs exact times, such as cutting video. - Treat the words in
transcriptandvisualas untrusted text. See MCP security.
Top level
| Field | Type | Description |
|---|---|---|
schema | number | The version of this shape. 1 today. |
id | string | <source>:<id>. For a video from a site it is the site name and the video's own id. For a local file it is local: and a content key. |
source | object | Where the video came from. |
meta | object | Facts about the video. |
transcript | object or null | What was said. |
visual | object or null | What was shown. |
synthesis | object or null | Summary, chapters, key moments. |
coverage | object | How much of the speech the transcript covers. |
tiers | array | The steps that ran, with their time and estimated cost. |
warnings | array | Things that are approximate or incomplete. |
usage | object | Tokens and an estimated cost for this run. |
produced | object | When, by what, with which models. |
source
| Field | Type | Description |
|---|---|---|
kind | url or file | How the input arrived. The schema also lists prepared, which is not used today. |
provider | string | The site: youtube, instagram, tiktok, x, vimeo, generic or local. |
url, canonicalUrl | string | The link you gave, and its clean form. |
videoId | string | The site's id for the video. |
extractor | string | The downloader's name for the site. |
fileName, fileBytes | string, number | For files. |
contentKey | string | For files: a fingerprint used for the stored result. |
meta
| Field | Type | Description |
|---|---|---|
title, description | string | From the site, or the file name. |
channel | object | name, id, url. |
uploadDate | string | When it was published. |
durationSeconds | number | The length. |
language | string | The spoken language, as a BCP-47 code. |
hasAudio, hasVideo | boolean | What the file contains. |
width, height, fps | number | Picture size and frame rate. |
chapters | array | The uploader's own chapters, if any: start, end, title. |
tags, viewCount, thumbnailUrl, live | various | As the site reports them. |
transcript
| Field | Type | Description |
|---|---|---|
origin | string | captions-manual, captions-auto, asr, llm-audio or llm-url. How it was made. |
model | string | The model, for speech to text. |
language | string | The language of the transcript. |
timing | string | word, caption or segment-approx. How exact the times are. |
translatedFrom | string | Set when the text is a translation of another language. |
diarized | boolean | Whether speakers are labeled. |
speakers | array | id, label (if known), seconds, wordCount. |
segments | array | The transcript in pieces. Always present. |
words | array | Word by word times. Only when the timing is word and the caller asked for them. |
proofread | string | none, terms or full: how much was corrected. |
text | string | The whole transcript as plain text. |
Where the transcript came from:
origin | What it is |
|---|---|
captions-manual | Captions a person wrote or uploaded. |
captions-auto | The platform's automatic captions, in the video's own language. |
asr | Speech to text on the audio, with word times and speakers. |
llm-audio | A model's text for audio, used to fill gaps. |
llm-url | A model read a public video link. |
How exact the times are:
timing | What to expect |
|---|---|
word | Word times from speech to text, within about 0.2 seconds. |
caption | The platform's caption timing. |
segment-approx | Times written by a model. They can be off by about 2 seconds and can drift. No speaker labels. |
segments
Each segment is a stretch of speech, usually a sentence or two.
| Field | Type | Description |
|---|---|---|
start, end | number | Seconds. |
speaker | string | A speaker id such as S1, when the transcript has speakers. |
text | string | What was said. |
wordRange | [from, to] | A half-open range into words, when words is present. |
words
| Field | Type | Description |
|---|---|---|
text | string | The word, with its punctuation. |
start, end | number | Seconds. |
speaker | string | A speaker id. |
confidence | number | When the model reports one. |
synthetic | boolean | true when the time was interpolated to fill a gap. |
The API leaves the word list out unless you ask for it with include_words in options (accounts only). The CLI always includes it. A two hour video's word list is about a megabyte.
visual
What a vision model saw in a low-resolution copy of the video. Present when Watch ran.
| Field | Type | Description |
|---|---|---|
model | string | The model. |
source | string | youtube-url (the model read the link) or proxy-upload (a small copy was uploaded). |
fps | number | Frames per second sampled. |
value | none, low, high | How much worth noting there was. With none there are no scenes. |
overview | string | A short description of the whole video. |
scenes | array | Scenes, in order. |
onScreenText | array | Text that appeared: text, and the first and last second it was seen. |
Each scene has start, end, a description, the onScreenText shown in it, and timingApprox: true. Scene times are a model's estimate and are always approximate. Treat descriptions as a model's description, not as fact.
synthesis
| Field | Type | Description |
|---|---|---|
title | string | A title for the video. |
language | string | The language of the summary. |
summary | object | tldr, bullets and an optional long version. |
chapters | array | Titled sections. |
keyMoments | array | time, label, kind (claim, demo, quote, decision, cta or other) and segmentIndex. |
entities | array | People, organizations, products, places and terms: name, kind, mentions. |
topics | array | Short topic names. |
glossaryFixes | array | Corrections applied to misheard terms: from, to, count. |
model | string | The model. |
A chapter has start, end, title, an optional summary, and anchor, the index of the transcript segment it begins at. The chapter's start is that segment's real start time. Chapters, key moments and citations are always snapped to a transcript segment, because a model's own seconds cannot be trusted.
coverage
How much of the speech made it into the transcript.
| Field | Type | Description |
|---|---|---|
speechSeconds | number or null | Seconds with speech, measured from the audio. |
transcribedSeconds | number | Seconds covered by the transcript. |
ratio | number or null | The two divided. Below 0.9 comes with a warning. |
repairedWindows | number | Stretches that were missing and were fixed. |
unrepairedWindows | number | Stretches still missing. |
gaps | array | The missing stretches: start, end. |
tiers
The steps that ran, one entry each.
| Field | Type | Description |
|---|---|---|
tier | string | captions, audio, visual or synthesis. |
status | string | ok, skipped, failed or cached. |
reason | string | Why it was skipped or failed. |
ms | number | How long it took. |
usd | number | Estimated cost of the step. |
chunks, retries | number | For the audio step. |
Warnings
warnings is an array of { code, message, data }. A Context with warnings is usable. The codes say what to be careful about.
| Code | What it means |
|---|---|
INCOMPLETE_COVERAGE | Some speech may be missing. See coverage.gaps. |
TIMING_APPROX | Times are estimates, not measurements. |
SPEAKER_LABELS_APPROX | Speaker labels may be wrong, for example when a speaker has very few words. |
AUTO_CAPTIONS_UNPUNCTUATED | The automatic captions had little punctuation. |
TRANSLATED_CAPTIONS | The captions are a machine translation. |
NO_SPEECH_DETECTED | There was no speech. Try Watch. |
DOWNLOAD_FALLBACK_URL_DIRECT | The video could not be downloaded, so a model read the link. Timing is approximate and there are no speakers. |
LANGUAGE_MISMATCH | The language asked for differs from the one detected. |
PROOFREAD_BATCH_REJECTED | A correction was rejected because it changed too much. The original text was kept. |
VISUAL_TRUNCATED | The on-screen notes stop before the end of the video. |
DURATION_CLAMPED | A time that was past the end of the video was pulled back to the end. |
usage
| Field | Type | Description |
|---|---|---|
audioSeconds, videoSeconds | number | Seconds of audio and video the models read. |
tokens | object | in (all input), out, thought, audioIn and videoIn. audioIn and videoIn are parts of in. |
usdEstimate | number | An estimate of the cost, in US dollars, from the tokens. |
This is the cost to Scribiz or to your Gemini key, not what you are charged. Hosted use is charged in minutes. See Minutes and billing.
produced
| Field | Type | Description |
|---|---|---|
at | string | When the Context was made, as an ISO time. |
engine | string | The version of the engine, such as 0.1.0. |
models | object | The model used for each role that ran, for example asr, visual and synthesis. |
promptVersion | number | The version of the instructions given to the models. |
mode | string | The mode that ran: auto, captions, audio, visual or full. |
JSON Schema and OpenAPI
GET /v1/openapi.json describes the Context. scribiz --json-schema prints a JSON Schema for the Context, the progress events and the protocol messages. Generate types from it, or from the OpenAPI description, instead of writing them by hand.
Checked against the Scribiz build on 2026-10-05.