Skip to content
Get transcript

Start

Speech models

Pick what reads your audio: Scribiz, your own Gemini or OpenRouter key, or Whisper on your own computer for free, offline and with no account.

View Markdown

Note

Scribiz as the model, on the web, in the API and in the chats, is live. Choosing another model is in the command-line tool 0.1.5 (npm install -g scribiz@latest) and in Mac app 0.1.5, which you can download now. On an older version, a signed-in CLI reads with Google's model and your own Gemini key still works.

Scribiz reads the audio of a video with a speech model. You can leave that choice to Scribiz, bring your own key, or run a model on your own computer. It is one list, one setting, in the command line and in the Mac app.

The list

ModelIdRunsYou needIt costs
Scribiz (the default)scribizOn Scribiz serversA Scribiz sign-inYour Scribiz minutes
GeminigeminiAt Google, with your keyA Gemini keyWhat Google charges, about $0.006 a minute
Scribe v2openrouter:elevenlabs/scribe-v2At OpenRouter, with your keyAn OpenRouter keyAbout $0.004 a minute
MAI-Transcribe 2openrouter:microsoft/mai-transcribe-2At OpenRouter, with your keyAn OpenRouter keyAbout $0.006 a minute
Grok STTopenrouter:x-ai/grok-stt-1.0At OpenRouter, with your keyAn OpenRouter keyAbout $0.002 a minute
Local Fastlocal:baseOn your computerA one-time download of 148 MBNothing
Local Balancedlocal:smallOn your computerA one-time download of 488 MBNothing
Local Bestlocal:turboOn your computerA one-time download of 574 MBNothing

scribiz models prints this list in your terminal with speed and accuracy bars and what each one still needs. In the Mac app it is Settings, Models.

What Scribiz is

The model Scribiz runs for you. Today that is ElevenLabs Scribe v2, reached through OpenRouter, with Google's Gemini behind it when Scribe cannot answer. Scribiz measured it against nine other models on short clips, meetings, quiet voice notes and 30 minute recordings. It had fewer wrong words than the Gemini model it replaced, read a long recording in about half the time, and labels who is speaking.

Counted the same way as before: a recording costs the same minutes whichever model reads it. See Minutes and billing.

Bring your own key

With your own key you pay the provider and Scribiz charges nothing.

  • Gemini. scribiz setup. Audio goes from your computer to Google.
  • OpenRouter. Make a key at openrouter.ai/keys, then scribiz models key and scribiz models use openrouter:elevenlabs/scribe-v2. Audio goes from your computer to OpenRouter, and Scribiz asks it to use only providers that do not keep or train on what they are sent. OPENROUTER_API_KEY in the environment works too. Any OpenRouter speech model that returns word times works: scribiz audio.mp3 --speech openrouter:<model>. The three in the list are the ones Scribiz measured.

Free, on your computer

Whisper runs on your computer through whisper.cpp. No account, no key, no upload, and it works offline.

Terminal
brew install whisper-cppscribiz models download basescribiz models use local:basescribiz talk.mp4

turbo (574 MB) is the one to use if you have the room: on English recordings it made about as many mistakes as the Gemini model Scribiz used until October 2026. base is instant and rough.

What a local model gives up, said plainly:

  • A transcript only. A summary, chapters, on-screen notes and the proofread need Gemini or a Scribiz sign-in. With neither, --detail full and --mode watch say so before anything runs.
  • No speaker labels. whisper.cpp does not tell voices apart.
  • Looser word times. Fewer than half of the word starts land within a fifth of a second of where Scribiz's model puts them. Fine to read, rough for subtitles.
  • More wrong words on hard audio. Meetings, quiet voices and languages other than English come out worse than with Scribiz.

In the Mac app, the Models page downloads a model with a progress bar and removes it again. whisper.cpp itself still has to be installed with Homebrew once, and the app gives a local model its compressed audio, so it makes somewhat more mistakes there than the command line does.

Choose one

You wantDo
The best result with nothing to set upStay on Scribiz
To pay Google or OpenRouter yourselfscribiz models use gemini or openrouter:<model>
Free, private, offlinescribiz models download turbo, scribiz models use local:turbo
One different model for one run--speech <id>
To see the choice in forcescribiz models or scribiz doctor

scribiz models use saves the choice in ~/.scribiz/config.json. --speech wins over SCRIBIZ_SPEECH, which wins over the saved choice, which wins over the default.

Where your audio goes

ModelAudio goes to
ScribizScribiz, then ElevenLabs through OpenRouter, or Google when that cannot answer
GeminiGoogle, with your key
OpenRouter modelsOpenRouter, with your key
Local modelsNowhere. It stays on your computer.

Summaries, chapters and the proofread are written by Gemini whichever model heard the audio. See Privacy and data.

How the bars were measured

The speed and accuracy bars come from our own runs, not from the vendors' pages. Wrong content words per 100, scored against what the other model families agree on (a human check on meeting recordings with published transcripts gave the same order). 108 short clips: public clips, meetings, and desk-microphone recordings of the person who built Scribiz, some of them turned down to a quiet level. Then 90 minutes of 30 minute recordings, cut into 10 minute pieces the way Scribiz sends them.

ModelShort clips30 minute recordingsSeconds for 10 minutes
Scribiz (Scribe v2)1.91.99
Gemini3.32.818
Local Best (turbo)3.73.020 to 26 on an M-series Mac
Local Balanced (small)5.83.722
Local Fast (base)8.56.210

Lower is better. The best local model is as accurate as the model Scribiz used until 9 October 2026, for nothing and with no account. It is behind Scribiz on meetings, quiet voices and other languages, has no speaker labels, and its word times are looser (under half of its words start within 0.2 seconds of the cloud models').

Not measured yet: audio longer than 90 minutes, several speakers over a long recording, other languages at length, and real phone voice notes (the quiet clips are simulated).

Loading the index