# Output formats

> SRT, VTT, TXT, Markdown, Context and JSON, with a real sample of each and what the options change.

Page: https://scribiz.com/docs/cli/formats

Pick a format with `--format`. The samples below are real output for a 19 second public video, `https://youtu.be/jNQXAC9IVRw`, with long lines trimmed. Where a sample needs speakers or on-screen notes, it comes from a 12 second recording of two voices. The two caption formats are built for players and editors. The others are built for reading, for notes, for agents and for code.

| Format | For | Built from |
| --- | --- | --- |
| `srt` | Editors and players | Caption-length cues |
| `vtt` | Web video | The same cues, with dots in the timestamps |
| `txt` | Reading and copying | Paragraphs, no timestamps |
| `md` | Notes | Title, summary and paragraphs with start times |
| `context` | Agents | The whole Context as Markdown: front matter, summary, key moments and the transcript with on-screen notes |
| `json` | Code | The whole Context object |

SRT, VTT and TXT only need the transcript, so `scribiz <input>` builds just that and writes no summary. The other formats need the whole Context.

## SRT

```srt title="zoo.srt"
1
00:00:01,200 --> 00:00:04,259
All right, so here we are,
in front of the elephants

2
00:00:05,318 --> 00:00:07,974
the cool thing about these guys
is that they have really...

3
00:00:07,974 --> 00:00:12,616
really really long trunks

4
00:00:12,616 --> 00:00:14,367
and that's cool

5
00:00:14,421 --> 00:00:15,733
(baaaaaaaaaaahhh!!)

6
00:00:16,881 --> 00:00:19,352
and that's pretty much all there is to say
```

Cue numbers start at 1. Timestamps use a comma before the milliseconds. A cue wraps to two lines of at most 42 characters each.

With `--speakers`, a cue starts with the speaker:

```srt title="speech.srt, with --speakers"
1
00:00:00,300 --> 00:00:02,418
[S1] Welcome to the Scribus fixture test.

2
00:00:02,600 --> 00:00:03,800
[S1] My name is Samantha.

3
00:00:04,500 --> 00:00:05,800
[S2] Thank you, Samantha.
```

Without it, SRT and VTT leave the speakers out so the file drops straight into an editor. Use `--offset 01:00:00` to shift every time, for a timeline that starts at one hour.

A video with no speech has no cues, so its SRT is empty.

### How cues are cut

A cue ends at a sentence end, at a pause of 0.75 seconds or more, after 8 words, or after 3.5 seconds. A comma ends a cue once it holds four words or two seconds. A speaker change always starts a new cue. A site name and its ending, like `picspot` and `.co`, stay together. A cue is at most 84 characters. A cue shorter than 0.7 seconds is extended into the pause that follows it, to about a second when there is room. A cue that would be read too fast is held on screen longer, up to two seconds past the end of the speech.

Cue timings come from the word times of the transcript. They are never guessed. When a transcript has no word times (captions, or a model reading a link), each cue follows the platform's own caption timing, and a caption that has to be split is cut at punctuation and its time divided in proportion to the length of each piece.

## VTT

```vtt title="zoo.vtt"
WEBVTT

00:00:01.200 --> 00:00:04.259
All right, so here we are,
in front of the elephants

00:00:05.318 --> 00:00:07.974
the cool thing about these guys
is that they have really...
```

Same cues as SRT, a `WEBVTT` header and a dot before the milliseconds. With `--speakers`, each cue carries a voice tag:

```vtt title="speech.vtt, with --speakers"
00:00:00.300 --> 00:00:02.418
<v S1>Welcome to the Scribus fixture test.</v>
```

## TXT

```text title="zoo.txt"
All right, so here we are, in front of the elephants

the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say
```

Paragraphs, not cues, with no timestamps. A paragraph breaks on a speaker change, a pause of 1.5 seconds or more, or after about five sentences. When the transcript has two or more speakers, each paragraph starts with the speaker:

```text title="speech.txt"
[S1]: Welcome to the Scribus fixture test. My name is Samantha.

[S2]: Thank you, Samantha. I am Daniel and the quick brown fox jumps over the lazy dog.

[S1]: The magic number is 42.
```

`--no-speakers` removes the labels. `--speakers` adds them for a transcript with one voice.

## MD

```markdown title="zoo.md"
# Me at the zoo

jawed · 00:19 · 2005-04-24 · https://www.youtube.com/watch?v=jNQXAC9IVRw · language en · transcript: platform captions (written by a person), caption timing from the platform

## Summary

The speaker stands in front of elephants and points out their remarkably long trunks before concluding the brief remark.

- The speaker arrives in front of the elephants at the zoo.
- Elephants have remarkably long trunks, which the speaker notes is cool.
- The speaker concludes that there is pretty much nothing more to say.

The speaker is recorded standing directly in front of the elephants. ...

## Transcript

[00:01] All right, so here we are, in front of the elephants

[00:05] the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to say
```

Markdown for people: a title, a line saying where the transcript came from, the summary with a longer write-up, then the transcript as paragraphs with their start times. Chapter titles appear as headings between paragraphs. Speakers show as `**S1:**` when there are two or more. Text the model wrote is escaped, so it cannot turn into Markdown structure.

## Context

The `context` format is the one to give an agent or paste into a chat. It is Markdown with front matter that says where the text came from, a summary, chapters when there are any, key moments, and the transcript with what was on screen written in at the right moments.

```markdown title="talking.context.md"
---
title: "Scribus Fixture Test"
source: "talking.mp4"
duration: "00:12"
duration_seconds: 12
language: "en"
transcript: "speech recognition, word-level timing"
speakers: 2
on_screen: "described from the picture by a vision model, times are approximate"
topics: ["fixture test", "audio testing", "test phrases"]
transcript_coverage: 1
produced: "2026-10-04T06:05:33.450Z by 0.1.0"
---

# Scribus Fixture Test

## Summary

Samantha and Daniel introduce the Scribus fixture test with sample phrases and a magic number. A blue title card displays on screen throughout the brief test.

- Samantha introduces the Scribus fixture test.
- Daniel recites the pangram about the quick brown fox jumping over the lazy dog.
- Samantha concludes by stating that the magic number is 42.

## Key moments

- [00:00] Samantha welcomes viewers to the fixture test
- [00:04] Daniel recites a standard test pangram (quote)
- [00:10] Samantha announces the magic number is 42 (claim)

## Transcript

[00:00] ON SCREEN: A blue title card showing the title 'SCRIBIZ FIXTURE 42' with a timer counting up at the bottom. Text: "SCRIBIZ FIXTURE 42"; "T+00s"

[00:00] **S1:** Welcome to the Scribus fixture test. My name is Samantha.

[00:04] **S2:** Thank you, Samantha. I am Daniel, and the quick brown fox jumps over the lazy dog.

[00:10] **S1:** The magic number is 42.
```

The front matter records the transcript's origin and timing, the transcript's coverage, and any warning codes, so an agent knows how far to trust it. `ON SCREEN` lines are descriptions from a vision model and their times are approximate. Treat them as notes about the picture, not as quotes. A video that was only listened to has no `ON SCREEN` lines.

A video with no speech lists its on-screen notes under an `On screen` heading instead, and the front matter says `transcript: "none"`. A video with chapters gets a `Chapters` list of `- 00:00 Title` lines before the key moments.

## JSON

```json title="zoo.json, abridged"
{
  "schema": 1,
  "id": "youtube:jNQXAC9IVRw",
  "source": {
    "kind": "url",
    "provider": "youtube",
    "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw",
    "videoId": "jNQXAC9IVRw"
  },
  "meta": {
    "title": "Me at the zoo",
    "durationSeconds": 19,
    "hasAudio": true,
    "hasVideo": true,
    "channel": { "name": "jawed" },
    "uploadDate": "2005-04-24"
  },
  "transcript": {
    "origin": "captions-manual",
    "language": "en",
    "timing": "caption",
    "diarized": false,
    "speakers": [],
    "segments": [
      { "start": 1.2, "end": 3.36, "text": "All right, so here we are, in front of the elephants" }
    ]
  },
  "visual": null,
  "synthesis": { "title": "Observing Elephants at the Zoo", "summary": { "tldr": "The speaker stands in front of elephants and points out that they have very long trunks." } },
  "warnings": []
}
```

This is the abridged Context object. [The Context object](https://scribiz.com/docs/api/context-object.md) lists every field. From the CLI, the JSON holds word times too, when the transcript has them. `--format json` writes it with indentation. `--json --out-file` writes it compact.

## Converting later

You do not need to run a video twice to get a second format. Save the JSON, then render it:

```bash
scribiz context ./talk.mp4 --format json -o talk.json
scribiz format talk.json --format srt -o talk.srt
scribiz format talk.json --format context -o talk.context.md
```

`scribiz format` needs no credential and no network. It also reads the file a `--json --out-file` run writes.

---

Checked against the Scribiz build on 2026-10-05.
