MCP
MCP limits and cost
What each tool call costs in minutes, how long a call can take, the limits on the remote server, and what the errors mean.
On this page
What a call costs
The tools use the same minutes as everywhere else. The first call on a video does the work and pays for it. Later calls on the same video read the stored result and cost nothing, except a summary Scribiz has not written for it yet, which uses a tenth of the video's length.
| Tool | Minutes |
|---|---|
get_video_context on a new video | By what runs, per minute of video: captions 0.1, Listen 1, Watch 1, Both 2 |
get_transcript | 0 once the video is processed. Otherwise the same as above |
search_video | 0 once the video is processed. Otherwise the same as above |
ask_video | 0.1 per question, on top of reading the video the first time. With a key, a question about something shown can run Watch, once, and watch: true runs it for any question. Without a key: 5 questions a day, answered from the words only (the picture is never looked at). A question still works when the day's 10 minutes are used up, for a video with captions or one already processed |
get_video_context or ask_video with watch: true (with a key) | Watch 1 per minute of video, once. The captions that come with it are free, and a Listen that had to run is added (Both is 2). 0 when the on-screen notes are already stored. See Watch, with a key |
| On-screen notes that are already stored (with a key) | 0, in every tool that reads a video |
get_job | 0 |
| Any call on a video someone already processed | 0, except a summary Scribiz has not written for it yet (a tenth of the video's length) |
Every result reports minutesUsed, and free: true when it came from stored results and cost nothing. A failed call is not charged for work that did not happen. If ask_video reads the video (Listen or Watch) and the question then fails, the read is charged, because it was done and is stored for everyone. A call that joins a run already going, or reuses one that just finished, is not charged again. See Minutes and billing.
The local server with GEMINI_API_KEY uses no Scribiz minutes. Google bills your key. It comes with the command-line tool (npm install -g scribiz, then claude mcp add scribiz -- scribiz mcp), needs ffmpeg and yt-dlp, and is tested on macOS only.
How long a call takes
Processing a long video takes minutes, and most clients give up on a tool call after about a minute. So a tool never blocks for long:
- It waits up to 45 seconds, sending progress updates if your client asked for them.
- If the video is not done, it returns
status: "processing"with ajob_idandretry_after_seconds. This is a normal result, not an error. - Calling again, or calling
get_job, picks up the run that is already going.
Typical timings: a video that was processed before answers at once. A public YouTube video with captions that Scribiz has not seen takes about 8 seconds. A short video read from its link takes a few seconds (about 4 seconds for a 19 second one). Listening to a longer video takes longer than one call waits, so expect a processing result and a second call.
Tokens
A tool result is text the agent has to read, so size matters.
| Result | Approximate size |
|---|---|
get_video_context with brief | Under 2,000 tokens. The 19 second example video is about 150. |
A page of get_transcript at the default max_chars | About 15,000 tokens |
search_video | A few hundred tokens |
| A two hour transcript read whole | More than 25,000 tokens |
Lower max_chars if your client warns about large tool results. Page with cursor.
Limits
| Limit | Value |
|---|---|
| Video length | Your plan's limit: 15 minutes without an account, 2 hours with a free account, 6 hours on Pro |
| Running videos per key | 2. A third call is RATE_LIMIT with a hint to call get_job. |
| Wait inside a call | 45 seconds |
| A finished run answers the same call again | For 15 minutes. With a key, a get_video_context or search_video result with no on-screen notes is run again for a YouTube link when notes have been stored since: it costs nothing |
| Local files | Local server only, and only from folders you allow |
| Remote server input | Links, not file paths |
| Without a key | Captions our server can read (30 lookups a day), stored results, and a model reading the link within the web tool's 10 minutes a day (videos up to 15 minutes), plus one daily limit shared by everyone without a key. Never the picture: no on-screen notes. See Without a key |
Longer videos than your plan allows are rejected before work starts.
Watch, with a key
A video that is mostly speech does not get the on-screen layer on its own. With a key you can have it, and you pay for it once.
| What | Cost |
|---|---|
| On-screen notes already stored, made by anyone | 0, in get_video_context, search_video and ask_video |
watch: true on a video whose notes are not stored | Watch: 1 minute for each minute of video. The captions that come with it are free. A Listen that had to run is added: Both is 2 |
watch: true again, or any other call, once the notes are stored | 0 |
Before anything is read, a watch: true call is quoted the way a job from the API is. It is checked against your plan's video length, the minutes you have left, and the minutes your other runs hold, and those minutes stay held while it runs. If it does not fit, it is refused at once with QUOTA or SOURCE_TOO_LONG, nothing is read, and nothing is charged. The message says how many minutes it needs. Call again without watch to leave the picture out.
A run that turns into a step you cannot pay for is stopped before that step starts. A layer that is skipped or fails is not charged. A link whose length could not be checked is read in a clip that ends at what you can still pay for; its notes are partial, they are not stored, and asking again reads again.
The picture is read once for a video. It is stored with the other layers, so the next call gets it free, with or without watch. A summary Scribiz has not written yet at that level of detail and in that language is the one thing that can still use a tenth of the video's length.
Calls at the same time for one video are taken one after the other, for one account. If you ask a question and for a summary of the same video together, the second waits for the first, finds the notes stored, and pays nothing for them. A wait that runs very long comes back as RATE_LIMIT, and the call can be repeated.
Without a key none of this applies. The picture is never looked at, a stored layer is not served, and watch is ignored with a note.
The local server (scribiz mcp, from the command-line tool, version 0.1.1 or newer) takes watch too, and serves on-screen notes that are already stored at no cost. What it costs goes to your credential: your Scribiz account's minutes after scribiz login, or Google's bill for your own Gemini key. See Connection modes.
Without a key
The server works without a key, inside a daily allowance. Days are UTC, and everything resets at 00:00 UTC.
| Allowance | Value |
|---|---|
| Reading | The same 10 minutes a day as the free web tool, for each connection. Reading is charged by the length of the video that was actually read: a 19 second video uses about 19 seconds. Captions use a tenth of that |
| Longest video | 15 minutes. A video known to be longer is refused before anything is used |
| Caption lookups | 30 a day for each connection |
| Questions | 5 a day, 0.1 minute each |
| Shared limit | One more daily limit of about $3 of reading, shared by everyone who connects without a key. It is separate from the website's free tools: neither can use up the other's. They share only your own 10 minutes a day |
What it does and does not do:
- It uses a video's captions when the server can read them, and serves a video Scribiz has already processed at once, however it was first read.
- When there are no captions the server can read, a model reads the video from its link. Its times are approximate (about 2 seconds either way), it has no speaker labels, and the result says so.
- It never looks at the picture, so there are no on-screen notes. A question about what was shown is answered from the words only.
- A summary written for a video that was already processed uses a tenth of its length (about 0.03 minutes for a 19 second video), and the result says so.
- If a read stops at its limit (a link whose length could not be checked in advance), the result says it is partial and how much was read. Repeating the call reads again and uses minutes again.
- One read runs at a time for each connection. A second call made at the same time waits for the first. If that takes longer than the wait, it comes back as
processingand the agent callsget_job. - When your 10 minutes, or the shared limit, are used up, captions and videos already processed still work. The error says what is used up and when it resets.
- For a video Scribiz has not processed, the link is passed to Google's model, which reads the public video. Nothing is uploaded from your machine.
Errors
A real failure is an error result, with isError: true. Its structured part has status: "failed" and an error object with the same codes the CLI and the API's jobs use, plus a message, a hint, a next_step for the agent, and retryable:
{ "status": "failed", "error": { "code": "URL_UNSUPPORTED", "message": "This address is not allowed: private or loopback address (169.254.169.254)", "retryable": false, "stage": "resolve", "hint": "Only public http(s) links to media or supported video sites are accepted here.", "next_step": "Pass a public http(s) link to a video. Local files and private addresses are not available on a hosted server." }}| Code | What it means | What the agent can do |
|---|---|---|
SOURCE_AUTH_REQUIRED | The site needs a login | Tell the user to download the file and upload it on the web, or to use the command-line tool with their browser's login (--cookies-from-browser) |
SOURCE_BLOCKED | The site refused a server | The same, or ask for another link |
SOURCE_UNAVAILABLE | Private, deleted or restricted | Ask for another link, or tell the user to download the file and upload it on scribiz.com |
SOURCE_TOO_LONG | Over your plan's limit. Without a key, a video known to be longer than 15 minutes | Ask for a shorter video |
URL_UNSUPPORTED | Not a link Scribiz can use, or a private address | Ask for a public link |
QUOTA, out of minutes | Without a key: the day's 10 minutes, the connection's spending limit, or the daily limit shared by everyone without a key. With a key, your minutes: a watch: true call needs a minute for each minute of video and is refused before anything is read when you have less | Tell the user. Without a key, captions and videos already processed still work (a key holder at zero minutes does not get this fallback). Without a key it resets at 00:00 UTC |
QUOTA, caption lookups | Without a key, the 30 caption lookups a day are used up on this connection | Tell the user. Videos already processed still work. It resets at 00:00 UTC |
QUOTA, questions | Without a key, the 5 questions a day are used up on this connection | Tell the user. Searching and reading videos still work. It resets at 00:00 UTC |
RATE_LIMIT | Too many calls, or too many videos running. Without a key, also minutes held by another read of yours that is still running: the error says what is held | Wait, then try again |
AUTH | The key was rejected | Tell the user to check SCRIBIZ_API_KEY |
TIMEOUT, NETWORK, SERVER | A temporary problem | Try again |
INVALID_ARGUMENTS | A bad argument, such as a time that does not parse, or an unknown job_id | Fix the call, or repeat the original call |
A video with no speech returns a result with the warning NO_SPEECH_DETECTED, not an error.
Every engine code is described in Errors and limits.
Checked against the Scribiz build on 2026-10-05.