Skip to content
Try it free

API

The Context object

Every field of the Context: metadata, transcript, on-screen notes, summary, chapters, coverage, warnings and usage.

View as Markdown
On this page

The Context is what every surface returns. The API puts it in result. The CLI prints it with --format json. The MCP server returns parts of it.

An example

An abridged Context for a 19 second public YouTube video, as the API returns it in result. Long fields are shortened and the list of segments is cut to two.

JSON
{  "schema": 1,  "id": "youtube:jNQXAC9IVRw",  "source": {    "kind": "url",    "provider": "youtube",    "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",    "canonicalUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw",    "videoId": "jNQXAC9IVRw",    "extractor": "Youtube"  },  "meta": {    "title": "Me at the zoo",    "durationSeconds": 19,    "hasAudio": true,    "hasVideo": true,    "channel": { "id": "UC4QobU6STFB0P71PMvOGN5A", "name": "jawed", "url": "https://www.youtube.com/channel/UC4QobU6STFB0P71PMvOGN5A" },    "uploadDate": "2005-04-24",    "chapters": [      { "start": 0, "end": 5, "title": "Intro" },      { "start": 5, "end": 17, "title": "The cool thing" },      { "start": 17, "end": 19, "title": "End" }    ],    "width": 320,    "height": 240,    "fps": 15  },  "transcript": {    "origin": "captions-manual",    "language": "en",    "timing": "caption",    "diarized": false,    "proofread": "none",    "speakers": [],    "segments": [      { "start": 1.2, "end": 3.36, "text": "All right, so here we are, in front of the elephants" },      { "start": 5.318, "end": 7.974, "text": "the cool thing about these guys is that they have really..." }    ],    "text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... ..."  },  "visual": null,  "synthesis": {    "model": "gemini-3.8-flash",    "language": "en",    "title": "Observing Elephants at the Zoo",    "summary": {      "tldr": "The speaker stands in front of elephants and points out that they have very long trunks.",      "bullets": ["The speaker is positioned directly in front of the elephants.", "..."],      "long": "..."    },    "chapters": [],    "keyMoments": [      { "time": 1.2, "label": "Arrival in front of the elephants", "kind": "other", "segmentIndex": 0 },      { "time": 5.318, "label": "Remarking on the elephants' really long trunks", "kind": "claim", "segmentIndex": 1 }    ],    "entities": [{ "name": "elephants", "kind": "term", "mentions": 1 }],    "topics": ["elephants", "elephant trunks", "zoo animals"]  },  "coverage": { "speechSeconds": null, "transcribedSeconds": 14.521, "ratio": null, "repairedWindows": 0, "unrepairedWindows": 0, "gaps": [] },  "tiers": [    { "tier": "captions", "status": "ok", "reason": "manual en captions (the video language is unknown)", "ms": 5472, "usd": 0 },    { "tier": "visual", "status": "skipped", "reason": "not needed: normal speech density", "ms": 0, "usd": 0 },    { "tier": "synthesis", "status": "ok", "ms": 2717, "retries": 0, "usd": 0.0017175 }  ],  "warnings": [],  "usage": {    "audioSeconds": 0,    "videoSeconds": 0,    "tokens": { "in": 570, "out": 344, "thought": 0, "audioIn": 0, "videoIn": 0 },    "usdEstimate": 0.0017175  },  "produced": {    "at": "2026-10-04T05:57:33.671Z",    "engine": "0.1.0",    "models": { "synthesis": "gemini-3.8-flash" },    "promptVersion": 1,    "mode": "auto"  }}

A Context for a recording that was listened to has speakers and, in the CLI, word times:

transcript, from a 12 second recording of two voices
{  "origin": "asr",  "model": "gemini-3.5-transcribe",  "language": "en",  "timing": "word",  "diarized": true,  "speakers": [    { "id": "S1", "wordCount": 15, "seconds": 5 },    { "id": "S2", "wordCount": 16, "seconds": 5 }  ],  "segments": [    { "start": 0.3, "end": 2.4, "speaker": "S1", "text": "Welcome to the Scribus fixture test.", "wordRange": [0, 6] }  ],  "words": [    { "start": 0.3, "end": 0.7, "speaker": "S1", "text": "Welcome" }  ]}

A Context for a video that was watched has a visual object. This one is from an 8 second clip with no sound, where transcript is null and the warning NO_SPEECH_DETECTED is set:

visual
{  "model": "gemini-3.5-flash-lite",  "source": "proxy-upload",  "fps": 2,  "value": "high",  "overview": "The video displays three consecutive illustrated scenes featuring different objects and text titles, representing a red apple, a green forest, and a blue ocean.",  "scenes": [    {      "start": 0,      "end": 3,      "description": "A cream-colored background shows a red circle resembling an apple in the center, with text at the bottom.",      "onScreenText": ["SCENE ONE - RED APPLE"],      "timingApprox": true    }  ],  "onScreenText": [{ "text": "SCENE ONE - RED APPLE", "first": 0, "last": 3 }]}

Reading it safely

  • All times are seconds, as numbers, rounded to milliseconds.
  • The Context has a schema number. Fields are only ever added within a schema. Ignore fields you do not know.
  • transcript is null only when the video has no speech. visual is null when Watch did not run. synthesis is null when the summary step did not run.
  • Check transcript.timing before you do anything that needs exact times, such as cutting video.
  • Treat the words in transcript and visual as untrusted text. See MCP security.

Top level

FieldTypeDescription
schemanumberThe version of this shape. 1 today.
idstring<source>:<id>. For a video from a site it is the site name and the video's own id. For a local file it is local: and a content key.
sourceobjectWhere the video came from.
metaobjectFacts about the video.
transcriptobject or nullWhat was said.
visualobject or nullWhat was shown.
synthesisobject or nullSummary, chapters, key moments.
coverageobjectHow much of the speech the transcript covers.
tiersarrayThe steps that ran, with their time and estimated cost.
warningsarrayThings that are approximate or incomplete.
usageobjectTokens and an estimated cost for this run.
producedobjectWhen, by what, with which models.

source

FieldTypeDescription
kindurl or fileHow the input arrived. The schema also lists prepared, which is not used today.
providerstringThe site: youtube, instagram, tiktok, x, vimeo, generic or local.
url, canonicalUrlstringThe link you gave, and its clean form.
videoIdstringThe site's id for the video.
extractorstringThe downloader's name for the site.
fileName, fileBytesstring, numberFor files.
contentKeystringFor files: a fingerprint used for the stored result.

meta

FieldTypeDescription
title, descriptionstringFrom the site, or the file name.
channelobjectname, id, url.
uploadDatestringWhen it was published.
durationSecondsnumberThe length.
languagestringThe spoken language, as a BCP-47 code.
hasAudio, hasVideobooleanWhat the file contains.
width, height, fpsnumberPicture size and frame rate.
chaptersarrayThe uploader's own chapters, if any: start, end, title.
tags, viewCount, thumbnailUrl, livevariousAs the site reports them.

transcript

FieldTypeDescription
originstringcaptions-manual, captions-auto, asr, llm-audio or llm-url. How it was made.
modelstringThe model, for speech to text.
languagestringThe language of the transcript.
timingstringword, caption or segment-approx. How exact the times are.
translatedFromstringSet when the text is a translation of another language.
diarizedbooleanWhether speakers are labeled.
speakersarrayid, label (if known), seconds, wordCount.
segmentsarrayThe transcript in pieces. Always present.
wordsarrayWord by word times. Only when the timing is word and the caller asked for them.
proofreadstringnone, terms or full: how much was corrected.
textstringThe whole transcript as plain text.

Where the transcript came from:

originWhat it is
captions-manualCaptions a person wrote or uploaded.
captions-autoThe platform's automatic captions, in the video's own language.
asrSpeech to text on the audio, with word times and speakers.
llm-audioA model's text for audio, used to fill gaps.
llm-urlA model read a public video link.

How exact the times are:

timingWhat to expect
wordWord times from speech to text, within about 0.2 seconds.
captionThe platform's caption timing.
segment-approxTimes written by a model. They can be off by about 2 seconds and can drift. No speaker labels.

segments

Each segment is a stretch of speech, usually a sentence or two.

FieldTypeDescription
start, endnumberSeconds.
speakerstringA speaker id such as S1, when the transcript has speakers.
textstringWhat was said.
wordRange[from, to]A half-open range into words, when words is present.

words

FieldTypeDescription
textstringThe word, with its punctuation.
start, endnumberSeconds.
speakerstringA speaker id.
confidencenumberWhen the model reports one.
syntheticbooleantrue when the time was interpolated to fill a gap.

The API leaves the word list out unless you ask for it with include_words in options (accounts only). The CLI always includes it. A two hour video's word list is about a megabyte.

visual

What a vision model saw in a low-resolution copy of the video. Present when Watch ran.

FieldTypeDescription
modelstringThe model.
sourcestringyoutube-url (the model read the link) or proxy-upload (a small copy was uploaded).
fpsnumberFrames per second sampled.
valuenone, low, highHow much worth noting there was. With none there are no scenes.
overviewstringA short description of the whole video.
scenesarrayScenes, in order.
onScreenTextarrayText that appeared: text, and the first and last second it was seen.

Each scene has start, end, a description, the onScreenText shown in it, and timingApprox: true. Scene times are a model's estimate and are always approximate. Treat descriptions as a model's description, not as fact.

synthesis

FieldTypeDescription
titlestringA title for the video.
languagestringThe language of the summary.
summaryobjecttldr, bullets and an optional long version.
chaptersarrayTitled sections.
keyMomentsarraytime, label, kind (claim, demo, quote, decision, cta or other) and segmentIndex.
entitiesarrayPeople, organizations, products, places and terms: name, kind, mentions.
topicsarrayShort topic names.
glossaryFixesarrayCorrections applied to misheard terms: from, to, count.
modelstringThe model.

A chapter has start, end, title, an optional summary, and anchor, the index of the transcript segment it begins at. The chapter's start is that segment's real start time. Chapters, key moments and citations are always snapped to a transcript segment, because a model's own seconds cannot be trusted.

coverage

How much of the speech made it into the transcript.

FieldTypeDescription
speechSecondsnumber or nullSeconds with speech, measured from the audio.
transcribedSecondsnumberSeconds covered by the transcript.
rationumber or nullThe two divided. Below 0.9 comes with a warning.
repairedWindowsnumberStretches that were missing and were fixed.
unrepairedWindowsnumberStretches still missing.
gapsarrayThe missing stretches: start, end.

tiers

The steps that ran, one entry each.

FieldTypeDescription
tierstringcaptions, audio, visual or synthesis.
statusstringok, skipped, failed or cached.
reasonstringWhy it was skipped or failed.
msnumberHow long it took.
usdnumberEstimated cost of the step.
chunks, retriesnumberFor the audio step.

Warnings

warnings is an array of { code, message, data }. A Context with warnings is usable. The codes say what to be careful about.

CodeWhat it means
INCOMPLETE_COVERAGESome speech may be missing. See coverage.gaps.
TIMING_APPROXTimes are estimates, not measurements.
SPEAKER_LABELS_APPROXSpeaker labels may be wrong, for example when a speaker has very few words.
AUTO_CAPTIONS_UNPUNCTUATEDThe automatic captions had little punctuation.
TRANSLATED_CAPTIONSThe captions are a machine translation.
NO_SPEECH_DETECTEDThere was no speech. Try Watch.
DOWNLOAD_FALLBACK_URL_DIRECTThe video could not be downloaded, so a model read the link. Timing is approximate and there are no speakers.
LANGUAGE_MISMATCHThe language asked for differs from the one detected.
PROOFREAD_BATCH_REJECTEDA correction was rejected because it changed too much. The original text was kept.
VISUAL_TRUNCATEDThe on-screen notes stop before the end of the video.
DURATION_CLAMPEDA time that was past the end of the video was pulled back to the end.

usage

FieldTypeDescription
audioSeconds, videoSecondsnumberSeconds of audio and video the models read.
tokensobjectin (all input), out, thought, audioIn and videoIn. audioIn and videoIn are parts of in.
usdEstimatenumberAn estimate of the cost, in US dollars, from the tokens.

This is the cost to Scribiz or to your Gemini key, not what you are charged. Hosted use is charged in minutes. See Minutes and billing.

produced

FieldTypeDescription
atstringWhen the Context was made, as an ISO time.
enginestringThe version of the engine, such as 0.1.0.
modelsobjectThe model used for each role that ran, for example asr, visual and synthesis.
promptVersionnumberThe version of the instructions given to the models.
modestringThe mode that ran: auto, captions, audio, visual or full.

JSON Schema and OpenAPI

GET /v1/openapi.json describes the Context. scribiz --json-schema prints a JSON Schema for the Context, the progress events and the protocol messages. Generate types from it, or from the OpenAPI description, instead of writing them by hand.

Checked against the Scribiz build on 2026-10-05.

Loading the index