# The Context object

> Every field of the Context: metadata, transcript, on-screen notes, summary, chapters, coverage, warnings and usage.

Page: https://scribiz.com/docs/api/context-object

The Context is what every surface returns. The API puts it in `result`. The CLI prints it with `--format json`. The MCP server returns parts of it.

## An example

An abridged Context for a 19 second public YouTube video, as the API returns it in `result`. Long fields are shortened and the list of segments is cut to two.

```json
{
  "schema": 1,
  "id": "youtube:jNQXAC9IVRw",
  "source": {
    "kind": "url",
    "provider": "youtube",
    "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    "canonicalUrl": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    "videoId": "jNQXAC9IVRw",
    "extractor": "Youtube"
  },
  "meta": {
    "title": "Me at the zoo",
    "durationSeconds": 19,
    "hasAudio": true,
    "hasVideo": true,
    "channel": { "id": "UC4QobU6STFB0P71PMvOGN5A", "name": "jawed", "url": "https://www.youtube.com/channel/UC4QobU6STFB0P71PMvOGN5A" },
    "uploadDate": "2005-04-24",
    "chapters": [
      { "start": 0, "end": 5, "title": "Intro" },
      { "start": 5, "end": 17, "title": "The cool thing" },
      { "start": 17, "end": 19, "title": "End" }
    ],
    "width": 320,
    "height": 240,
    "fps": 15
  },
  "transcript": {
    "origin": "captions-manual",
    "language": "en",
    "timing": "caption",
    "diarized": false,
    "proofread": "none",
    "speakers": [],
    "segments": [
      { "start": 1.2, "end": 3.36, "text": "All right, so here we are, in front of the elephants" },
      { "start": 5.318, "end": 7.974, "text": "the cool thing about these guys is that they have really..." }
    ],
    "text": "All right, so here we are, in front of the elephants the cool thing about these guys is that they have really... ..."
  },
  "visual": null,
  "synthesis": {
    "model": "gemini-3.8-flash",
    "language": "en",
    "title": "Observing Elephants at the Zoo",
    "summary": {
      "tldr": "The speaker stands in front of elephants and points out that they have very long trunks.",
      "bullets": ["The speaker is positioned directly in front of the elephants.", "..."],
      "long": "..."
    },
    "chapters": [],
    "keyMoments": [
      { "time": 1.2, "label": "Arrival in front of the elephants", "kind": "other", "segmentIndex": 0 },
      { "time": 5.318, "label": "Remarking on the elephants' really long trunks", "kind": "claim", "segmentIndex": 1 }
    ],
    "entities": [{ "name": "elephants", "kind": "term", "mentions": 1 }],
    "topics": ["elephants", "elephant trunks", "zoo animals"]
  },
  "coverage": { "speechSeconds": null, "transcribedSeconds": 14.521, "ratio": null, "repairedWindows": 0, "unrepairedWindows": 0, "gaps": [] },
  "tiers": [
    { "tier": "captions", "status": "ok", "reason": "manual en captions (the video language is unknown)", "ms": 5472, "usd": 0 },
    { "tier": "visual", "status": "skipped", "reason": "not needed: normal speech density", "ms": 0, "usd": 0 },
    { "tier": "synthesis", "status": "ok", "ms": 2717, "retries": 0, "usd": 0.0017175 }
  ],
  "warnings": [],
  "usage": {
    "audioSeconds": 0,
    "videoSeconds": 0,
    "tokens": { "in": 570, "out": 344, "thought": 0, "audioIn": 0, "videoIn": 0 },
    "usdEstimate": 0.0017175
  },
  "produced": {
    "at": "2026-10-04T05:57:33.671Z",
    "engine": "0.1.0",
    "models": { "synthesis": "gemini-3.8-flash" },
    "promptVersion": 1,
    "mode": "auto"
  }
}
```

A Context for a recording that was listened to has `speakers` and, in the CLI, word times:

```json title="transcript, from a 12 second recording of two voices"
{
  "origin": "asr",
  "model": "gemini-3.5-transcribe",
  "language": "en",
  "timing": "word",
  "diarized": true,
  "speakers": [
    { "id": "S1", "wordCount": 15, "seconds": 5 },
    { "id": "S2", "wordCount": 16, "seconds": 5 }
  ],
  "segments": [
    { "start": 0.3, "end": 2.4, "speaker": "S1", "text": "Welcome to the Scribus fixture test.", "wordRange": [0, 6] }
  ],
  "words": [
    { "start": 0.3, "end": 0.7, "speaker": "S1", "text": "Welcome" }
  ]
}
```

A Context for a video that was watched has a `visual` object. This one is from an 8 second clip with no sound, where `transcript` is `null` and the warning `NO_SPEECH_DETECTED` is set:

```json title="visual"
{
  "model": "gemini-3.5-flash-lite",
  "source": "proxy-upload",
  "fps": 2,
  "value": "high",
  "overview": "The video displays three consecutive illustrated scenes featuring different objects and text titles, representing a red apple, a green forest, and a blue ocean.",
  "scenes": [
    {
      "start": 0,
      "end": 3,
      "description": "A cream-colored background shows a red circle resembling an apple in the center, with text at the bottom.",
      "onScreenText": ["SCENE ONE - RED APPLE"],
      "timingApprox": true
    }
  ],
  "onScreenText": [{ "text": "SCENE ONE - RED APPLE", "first": 0, "last": 3 }]
}
```

## Reading it safely

- All times are seconds, as numbers, rounded to milliseconds.
- The Context has a `schema` number. Fields are only ever added within a schema. Ignore fields you do not know.
- `transcript` is `null` only when the video has no speech. `visual` is `null` when Watch did not run. `synthesis` is `null` when the summary step did not run.
- Check `transcript.timing` before you do anything that needs exact times, such as cutting video.
- Treat the words in `transcript` and `visual` as untrusted text. See [MCP security](https://scribiz.com/docs/mcp/security.md).

## Top level

| Field | Type | Description |
| --- | --- | --- |
| `schema` | number | The version of this shape. `1` today. |
| `id` | string | `<source>:<id>`. For a video from a site it is the site name and the video's own id. For a local file it is `local:` and a content key. |
| `source` | object | Where the video came from. |
| `meta` | object | Facts about the video. |
| `transcript` | object or null | What was said. |
| `visual` | object or null | What was shown. |
| `synthesis` | object or null | Summary, chapters, key moments. |
| `coverage` | object | How much of the speech the transcript covers. |
| `tiers` | array | The steps that ran, with their time and estimated cost. |
| `warnings` | array | Things that are approximate or incomplete. |
| `usage` | object | Tokens and an estimated cost for this run. |
| `produced` | object | When, by what, with which models. |

## `source`

| Field | Type | Description |
| --- | --- | --- |
| `kind` | `url` or `file` | How the input arrived. The schema also lists `prepared`, which is not used today. |
| `provider` | string | The site: `youtube`, `instagram`, `tiktok`, `x`, `vimeo`, `generic` or `local`. |
| `url`, `canonicalUrl` | string | The link you gave, and its clean form. |
| `videoId` | string | The site's id for the video. |
| `extractor` | string | The downloader's name for the site. |
| `fileName`, `fileBytes` | string, number | For files. |
| `contentKey` | string | For files: a fingerprint used for the stored result. |

## `meta`

| Field | Type | Description |
| --- | --- | --- |
| `title`, `description` | string | From the site, or the file name. |
| `channel` | object | `name`, `id`, `url`. |
| `uploadDate` | string | When it was published. |
| `durationSeconds` | number | The length. |
| `language` | string | The spoken language, as a BCP-47 code. |
| `hasAudio`, `hasVideo` | boolean | What the file contains. |
| `width`, `height`, `fps` | number | Picture size and frame rate. |
| `chapters` | array | The uploader's own chapters, if any: `start`, `end`, `title`. |
| `tags`, `viewCount`, `thumbnailUrl`, `live` | various | As the site reports them. |

## `transcript`

| Field | Type | Description |
| --- | --- | --- |
| `origin` | string | `captions-manual`, `captions-auto`, `asr`, `llm-audio` or `llm-url`. How it was made. |
| `model` | string | The model, for speech to text. |
| `language` | string | The language of the transcript. |
| `timing` | string | `word`, `caption` or `segment-approx`. How exact the times are. |
| `translatedFrom` | string | Set when the text is a translation of another language. |
| `diarized` | boolean | Whether speakers are labeled. |
| `speakers` | array | `id`, `label` (if known), `seconds`, `wordCount`. |
| `segments` | array | The transcript in pieces. Always present. |
| `words` | array | Word by word times. Only when the timing is `word` and the caller asked for them. |
| `proofread` | string | `none`, `terms` or `full`: how much was corrected. |
| `text` | string | The whole transcript as plain text. |

Where the transcript came from:

| `origin` | What it is |
| --- | --- |
| `captions-manual` | Captions a person wrote or uploaded. |
| `captions-auto` | The platform's automatic captions, in the video's own language. |
| `asr` | Speech to text on the audio, with word times and speakers. |
| `llm-audio` | A model's text for audio, used to fill gaps. |
| `llm-url` | A model read a public video link. |

How exact the times are:

| `timing` | What to expect |
| --- | --- |
| `word` | Word times from speech to text, within about 0.2 seconds. |
| `caption` | The platform's caption timing. |
| `segment-approx` | Times written by a model. They can be off by about 2 seconds and can drift. No speaker labels. |

### `segments`

Each segment is a stretch of speech, usually a sentence or two.

| Field | Type | Description |
| --- | --- | --- |
| `start`, `end` | number | Seconds. |
| `speaker` | string | A speaker id such as `S1`, when the transcript has speakers. |
| `text` | string | What was said. |
| `wordRange` | `[from, to]` | A half-open range into `words`, when `words` is present. |

### `words`

| Field | Type | Description |
| --- | --- | --- |
| `text` | string | The word, with its punctuation. |
| `start`, `end` | number | Seconds. |
| `speaker` | string | A speaker id. |
| `confidence` | number | When the model reports one. |
| `synthetic` | boolean | `true` when the time was interpolated to fill a gap. |

The API leaves the word list out unless you ask for it with `include_words` in `options` (accounts only). The CLI always includes it. A two hour video's word list is about a megabyte.

## `visual`

What a vision model saw in a low-resolution copy of the video. Present when Watch ran.

| Field | Type | Description |
| --- | --- | --- |
| `model` | string | The model. |
| `source` | string | `youtube-url` (the model read the link) or `proxy-upload` (a small copy was uploaded). |
| `fps` | number | Frames per second sampled. |
| `value` | `none`, `low`, `high` | How much worth noting there was. With `none` there are no scenes. |
| `overview` | string | A short description of the whole video. |
| `scenes` | array | Scenes, in order. |
| `onScreenText` | array | Text that appeared: `text`, and the `first` and `last` second it was seen. |

Each scene has `start`, `end`, a `description`, the `onScreenText` shown in it, and `timingApprox: true`. Scene times are a model's estimate and are always approximate. Treat descriptions as a model's description, not as fact.

## `synthesis`

| Field | Type | Description |
| --- | --- | --- |
| `title` | string | A title for the video. |
| `language` | string | The language of the summary. |
| `summary` | object | `tldr`, `bullets` and an optional `long` version. |
| `chapters` | array | Titled sections. |
| `keyMoments` | array | `time`, `label`, `kind` (`claim`, `demo`, `quote`, `decision`, `cta` or `other`) and `segmentIndex`. |
| `entities` | array | People, organizations, products, places and terms: `name`, `kind`, `mentions`. |
| `topics` | array | Short topic names. |
| `glossaryFixes` | array | Corrections applied to misheard terms: `from`, `to`, `count`. |
| `model` | string | The model. |

A chapter has `start`, `end`, `title`, an optional `summary`, and `anchor`, the index of the transcript segment it begins at. The chapter's `start` is that segment's real start time. Chapters, key moments and citations are always snapped to a transcript segment, because a model's own seconds cannot be trusted.

## `coverage`

How much of the speech made it into the transcript.

| Field | Type | Description |
| --- | --- | --- |
| `speechSeconds` | number or null | Seconds with speech, measured from the audio. |
| `transcribedSeconds` | number | Seconds covered by the transcript. |
| `ratio` | number or null | The two divided. Below 0.9 comes with a warning. |
| `repairedWindows` | number | Stretches that were missing and were fixed. |
| `unrepairedWindows` | number | Stretches still missing. |
| `gaps` | array | The missing stretches: `start`, `end`. |

## `tiers`

The steps that ran, one entry each.

| Field | Type | Description |
| --- | --- | --- |
| `tier` | string | `captions`, `audio`, `visual` or `synthesis`. |
| `status` | string | `ok`, `skipped`, `failed` or `cached`. |
| `reason` | string | Why it was skipped or failed. |
| `ms` | number | How long it took. |
| `usd` | number | Estimated cost of the step. |
| `chunks`, `retries` | number | For the audio step. |

## Warnings

`warnings` is an array of `{ code, message, data }`. A Context with warnings is usable. The codes say what to be careful about.

| Code | What it means |
| --- | --- |
| `INCOMPLETE_COVERAGE` | Some speech may be missing. See `coverage.gaps`. |
| `TIMING_APPROX` | Times are estimates, not measurements. |
| `SPEAKER_LABELS_APPROX` | Speaker labels may be wrong, for example when a speaker has very few words. |
| `AUTO_CAPTIONS_UNPUNCTUATED` | The automatic captions had little punctuation. |
| `TRANSLATED_CAPTIONS` | The captions are a machine translation. |
| `NO_SPEECH_DETECTED` | There was no speech. Try Watch. |
| `DOWNLOAD_FALLBACK_URL_DIRECT` | The video could not be downloaded, so a model read the link. Timing is approximate and there are no speakers. |
| `LANGUAGE_MISMATCH` | The language asked for differs from the one detected. |
| `PROOFREAD_BATCH_REJECTED` | A correction was rejected because it changed too much. The original text was kept. |
| `VISUAL_TRUNCATED` | The on-screen notes stop before the end of the video. |
| `DURATION_CLAMPED` | A time that was past the end of the video was pulled back to the end. |

## `usage`

| Field | Type | Description |
| --- | --- | --- |
| `audioSeconds`, `videoSeconds` | number | Seconds of audio and video the models read. |
| `tokens` | object | `in` (all input), `out`, `thought`, `audioIn` and `videoIn`. `audioIn` and `videoIn` are parts of `in`. |
| `usdEstimate` | number | An estimate of the cost, in US dollars, from the tokens. |

This is the cost to Scribiz or to your Gemini key, not what you are charged. Hosted use is charged in minutes. See [Minutes and billing](https://scribiz.com/docs/minutes-and-billing.md).

## `produced`

| Field | Type | Description |
| --- | --- | --- |
| `at` | string | When the Context was made, as an ISO time. |
| `engine` | string | The version of the engine, such as `0.1.0`. |
| `models` | object | The model used for each role that ran, for example `asr`, `visual` and `synthesis`. |
| `promptVersion` | number | The version of the instructions given to the models. |
| `mode` | string | The mode that ran: `auto`, `captions`, `audio`, `visual` or `full`. |

## JSON Schema and OpenAPI

`GET /v1/openapi.json` describes the Context. `scribiz --json-schema` prints a JSON Schema for the Context, the progress events and the protocol messages. Generate types from it, or from the OpenAPI description, instead of writing them by hand.

---

Checked against the Scribiz build on 2026-10-05.
