CLI
Output formats
SRT, VTT, TXT, Markdown, Context and JSON, with a real sample of each and what the options change.
On this page
Pick a format with --format. The samples below are real output for a 19 second public video, https://youtu.be/jNQXAC9IVRw, with long lines trimmed. Where a sample needs speakers or on-screen notes, it comes from a 12 second recording of two voices. The two caption formats are built for players and editors. The others are built for reading, for notes, for agents and for code.
| Format | For | Built from |
|---|---|---|
srt | Editors and players | Caption-length cues |
vtt | Web video | The same cues, with dots in the timestamps |
txt | Reading and copying | Paragraphs, no timestamps |
md | Notes | Title, summary and paragraphs with start times |
context | Agents | The whole Context as Markdown: front matter, summary, key moments and the transcript with on-screen notes |
json | Code | The whole Context object |
SRT, VTT and TXT only need the transcript, so scribiz <input> builds just that and writes no summary. The other formats need the whole Context.
SRT
100:00:01,200 --> 00:00:04,259All right, so here we are,in front of the elephants200:00:05,318 --> 00:00:07,974the cool thing about these guysis that they have really...300:00:07,974 --> 00:00:12,616really really long trunks400:00:12,616 --> 00:00:14,367and that's cool500:00:14,421 --> 00:00:15,733(baaaaaaaaaaahhh!!)600:00:16,881 --> 00:00:19,352and that's pretty much all there is to sayCue numbers start at 1. Timestamps use a comma before the milliseconds. A cue wraps to two lines of at most 42 characters each.
With --speakers, a cue starts with the speaker:
100:00:00,300 --> 00:00:02,418[S1] Welcome to the Scribus fixture test.200:00:02,600 --> 00:00:03,800[S1] My name is Samantha.300:00:04,500 --> 00:00:05,800[S2] Thank you, Samantha.Without it, SRT and VTT leave the speakers out so the file drops straight into an editor. Use --offset 01:00:00 to shift every time, for a timeline that starts at one hour.
A video with no speech has no cues, so its SRT is empty.
How cues are cut
A cue ends at a sentence end, at a pause of 0.75 seconds or more, after 8 words, or after 3.5 seconds. A comma ends a cue once it holds four words or two seconds. A speaker change always starts a new cue. A site name and its ending, like picspot and .co, stay together. A cue is at most 84 characters. A cue shorter than 0.7 seconds is extended into the pause that follows it, to about a second when there is room. A cue that would be read too fast is held on screen longer, up to two seconds past the end of the speech.
Cue timings come from the word times of the transcript. They are never guessed. When a transcript has no word times (captions, or a model reading a link), each cue follows the platform's own caption timing, and a caption that has to be split is cut at punctuation and its time divided in proportion to the length of each piece.
VTT
WEBVTT00:00:01.200 --> 00:00:04.259All right, so here we are,in front of the elephants00:00:05.318 --> 00:00:07.974the cool thing about these guysis that they have really...Same cues as SRT, a WEBVTT header and a dot before the milliseconds. With --speakers, each cue carries a voice tag:
00:00:00.300 --> 00:00:02.418<v S1>Welcome to the Scribus fixture test.</v>TXT
All right, so here we are, in front of the elephantsthe cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to sayParagraphs, not cues, with no timestamps. A paragraph breaks on a speaker change, a pause of 1.5 seconds or more, or after about five sentences. When the transcript has two or more speakers, each paragraph starts with the speaker:
[S1]: Welcome to the Scribus fixture test. My name is Samantha.[S2]: Thank you, Samantha. I am Daniel and the quick brown fox jumps over the lazy dog.[S1]: The magic number is 42.--no-speakers removes the labels. --speakers adds them for a transcript with one voice.
MD
# Me at the zoojawed · 00:19 · 2005-04-24 · https://www.youtube.com/watch?v=jNQXAC9IVRw · language en · transcript: platform captions (written by a person), caption timing from the platform## SummaryThe speaker stands in front of elephants and points out their remarkably long trunks before concluding the brief remark.- The speaker arrives in front of the elephants at the zoo.- Elephants have remarkably long trunks, which the speaker notes is cool.- The speaker concludes that there is pretty much nothing more to say.The speaker is recorded standing directly in front of the elephants. ...## Transcript[00:01] All right, so here we are, in front of the elephants[00:05] the cool thing about these guys is that they have really... really really long trunks and that's cool (baaaaaaaaaaahhh!!) and that's pretty much all there is to sayMarkdown for people: a title, a line saying where the transcript came from, the summary with a longer write-up, then the transcript as paragraphs with their start times. Chapter titles appear as headings between paragraphs. Speakers show as **S1:** when there are two or more. Text the model wrote is escaped, so it cannot turn into Markdown structure.
Context
The context format is the one to give an agent or paste into a chat. It is Markdown with front matter that says where the text came from, a summary, chapters when there are any, key moments, and the transcript with what was on screen written in at the right moments.
---title: "Scribus Fixture Test"source: "talking.mp4"duration: "00:12"duration_seconds: 12language: "en"transcript: "speech recognition, word-level timing"speakers: 2on_screen: "described from the picture by a vision model, times are approximate"topics: ["fixture test", "audio testing", "test phrases"]transcript_coverage: 1produced: "2026-10-04T06:05:33.450Z by 0.1.0"---# Scribus Fixture Test## SummarySamantha and Daniel introduce the Scribus fixture test with sample phrases and a magic number. A blue title card displays on screen throughout the brief test.- Samantha introduces the Scribus fixture test.- Daniel recites the pangram about the quick brown fox jumping over the lazy dog.- Samantha concludes by stating that the magic number is 42.## Key moments- [00:00] Samantha welcomes viewers to the fixture test- [00:04] Daniel recites a standard test pangram (quote)- [00:10] Samantha announces the magic number is 42 (claim)## Transcript[00:00] ON SCREEN: A blue title card showing the title 'SCRIBIZ FIXTURE 42' with a timer counting up at the bottom. Text: "SCRIBIZ FIXTURE 42"; "T+00s"[00:00] **S1:** Welcome to the Scribus fixture test. My name is Samantha.[00:04] **S2:** Thank you, Samantha. I am Daniel, and the quick brown fox jumps over the lazy dog.[00:10] **S1:** The magic number is 42.The front matter records the transcript's origin and timing, the transcript's coverage, and any warning codes, so an agent knows how far to trust it. ON SCREEN lines are descriptions from a vision model and their times are approximate. Treat them as notes about the picture, not as quotes. A video that was only listened to has no ON SCREEN lines.
A video with no speech lists its on-screen notes under an On screen heading instead, and the front matter says transcript: "none". A video with chapters gets a Chapters list of - 00:00 Title lines before the key moments.
JSON
{ "schema": 1, "id": "youtube:jNQXAC9IVRw", "source": { "kind": "url", "provider": "youtube", "url": "https://www.youtube.com/watch?v=jNQXAC9IVRw", "videoId": "jNQXAC9IVRw" }, "meta": { "title": "Me at the zoo", "durationSeconds": 19, "hasAudio": true, "hasVideo": true, "channel": { "name": "jawed" }, "uploadDate": "2005-04-24" }, "transcript": { "origin": "captions-manual", "language": "en", "timing": "caption", "diarized": false, "speakers": [], "segments": [ { "start": 1.2, "end": 3.36, "text": "All right, so here we are, in front of the elephants" } ] }, "visual": null, "synthesis": { "title": "Observing Elephants at the Zoo", "summary": { "tldr": "The speaker stands in front of elephants and points out that they have very long trunks." } }, "warnings": []}This is the abridged Context object. The Context object lists every field. From the CLI, the JSON holds word times too, when the transcript has them. --format json writes it with indentation. --json --out-file writes it compact.
Converting later
You do not need to run a video twice to get a second format. Save the JSON, then render it:
scribiz context ./talk.mp4 --format json -o talk.jsonscribiz format talk.json --format srt -o talk.srtscribiz format talk.json --format context -o talk.context.mdscribiz format needs no credential and no network. It also reads the file a --json --out-file run writes.
Checked against the Scribiz build on 2026-10-05.