# Speech models

> Pick what reads your audio: Scribiz, your own Gemini or OpenRouter key, or Whisper on your own computer for free, offline and with no account.

Page: https://scribiz.com/docs/models

> [!NOTE]
> Scribiz as the model, on the web, in the API and in the chats, is live. Choosing another model is in the command-line tool 0.1.5 (`npm install -g scribiz@latest`) and in Mac app 0.1.5, which you can download now. On an older version, a signed-in CLI reads with Google's model and your own Gemini key still works.

Scribiz reads the audio of a video with a speech model. You can leave that choice to Scribiz, bring your own key, or run a model on your own computer. It is one list, one setting, in the command line and in the Mac app.

## The list

| Model | Id | Runs | You need | It costs |
| --- | --- | --- | --- | --- |
| **Scribiz** (the default) | `scribiz` | On Scribiz servers | A Scribiz sign-in | Your Scribiz minutes |
| **Gemini** | `gemini` | At Google, with your key | A Gemini key | What Google charges, about $0.006 a minute |
| **Scribe v2** | `openrouter:elevenlabs/scribe-v2` | At OpenRouter, with your key | An OpenRouter key | About $0.004 a minute |
| **MAI-Transcribe 2** | `openrouter:microsoft/mai-transcribe-2` | At OpenRouter, with your key | An OpenRouter key | About $0.006 a minute |
| **Grok STT** | `openrouter:x-ai/grok-stt-1.0` | At OpenRouter, with your key | An OpenRouter key | About $0.002 a minute |
| **Local Fast** | `local:base` | On your computer | A one-time download of 148 MB | Nothing |
| **Local Balanced** | `local:small` | On your computer | A one-time download of 488 MB | Nothing |
| **Local Best** | `local:turbo` | On your computer | A one-time download of 574 MB | Nothing |

`scribiz models` prints this list in your terminal with speed and accuracy bars and what each one still needs. In the Mac app it is Settings, Models.

## What Scribiz is

The model Scribiz runs for you. Today that is ElevenLabs Scribe v2, reached through OpenRouter, with Google's Gemini behind it when Scribe cannot answer. Scribiz measured it against nine other models on short clips, meetings, quiet voice notes and 30 minute recordings. It had fewer wrong words than the Gemini model it replaced, read a long recording in about half the time, and labels who is speaking.

Counted the same way as before: a recording costs the same minutes whichever model reads it. See [Minutes and billing](https://scribiz.com/docs/minutes-and-billing.md).

## Bring your own key

With your own key you pay the provider and Scribiz charges nothing.

- **Gemini.** `scribiz setup`. Audio goes from your computer to Google.
- **OpenRouter.** Make a key at [openrouter.ai/keys](https://openrouter.ai/keys), then `scribiz models key` and `scribiz models use openrouter:elevenlabs/scribe-v2`. Audio goes from your computer to OpenRouter, and Scribiz asks it to use only providers that do not keep or train on what they are sent. `OPENROUTER_API_KEY` in the environment works too. Any OpenRouter speech model that returns word times works: `scribiz audio.mp3 --speech openrouter:<model>`. The three in the list are the ones Scribiz measured.

## Free, on your computer

Whisper runs on your computer through whisper.cpp. No account, no key, no upload, and it works offline.

```bash
brew install whisper-cpp
scribiz models download base
scribiz models use local:base
scribiz talk.mp4
```

`turbo` (574 MB) is the one to use if you have the room: on English recordings it made about as many mistakes as the Gemini model Scribiz used until October 2026. `base` is instant and rough.

What a local model gives up, said plainly:

- **A transcript only.** A summary, chapters, on-screen notes and the proofread need Gemini or a Scribiz sign-in. With neither, `--detail full` and `--mode watch` say so before anything runs.
- **No speaker labels.** whisper.cpp does not tell voices apart.
- **Looser word times.** Fewer than half of the word starts land within a fifth of a second of where Scribiz's model puts them. Fine to read, rough for subtitles.
- **More wrong words on hard audio.** Meetings, quiet voices and languages other than English come out worse than with Scribiz.

In the Mac app, the Models page downloads a model with a progress bar and removes it again. whisper.cpp itself still has to be installed with Homebrew once, and the app gives a local model its compressed audio, so it makes somewhat more mistakes there than the command line does.

## Choose one

| You want | Do |
| --- | --- |
| The best result with nothing to set up | Stay on **Scribiz** |
| To pay Google or OpenRouter yourself | `scribiz models use gemini` or `openrouter:<model>` |
| Free, private, offline | `scribiz models download turbo`, `scribiz models use local:turbo` |
| One different model for one run | `--speech <id>` |
| To see the choice in force | `scribiz models` or `scribiz doctor` |

`scribiz models use` saves the choice in `~/.scribiz/config.json`. `--speech` wins over `SCRIBIZ_SPEECH`, which wins over the saved choice, which wins over the default.

## Where your audio goes

| Model | Audio goes to |
| --- | --- |
| Scribiz | Scribiz, then ElevenLabs through OpenRouter, or Google when that cannot answer |
| Gemini | Google, with your key |
| OpenRouter models | OpenRouter, with your key |
| Local models | Nowhere. It stays on your computer. |

Summaries, chapters and the proofread are written by Gemini whichever model heard the audio. See [Privacy and data](https://scribiz.com/docs/privacy-and-data.md).

## How the bars were measured

The speed and accuracy bars come from our own runs, not from the vendors' pages. Wrong content words per 100, scored against what the other model families agree on (a human check on meeting recordings with published transcripts gave the same order). 108 short clips: public clips, meetings, and desk-microphone recordings of the person who built Scribiz, some of them turned down to a quiet level. Then 90 minutes of 30 minute recordings, cut into 10 minute pieces the way Scribiz sends them.

| Model | Short clips | 30 minute recordings | Seconds for 10 minutes |
| --- | --- | --- | --- |
| Scribiz (Scribe v2) | 1.9 | 1.9 | 9 |
| Gemini | 3.3 | 2.8 | 18 |
| Local Best (turbo) | 3.7 | 3.0 | 20 to 26 on an M-series Mac |
| Local Balanced (small) | 5.8 | 3.7 | 22 |
| Local Fast (base) | 8.5 | 6.2 | 10 |

Lower is better. The best local model is as accurate as the model Scribiz used until 9 October 2026, for nothing and with no account. It is behind Scribiz on meetings, quiet voices and other languages, has no speaker labels, and its word times are looser (under half of its words start within 0.2 seconds of the cloud models').

Not measured yet: audio longer than 90 minutes, several speakers over a long recording, other languages at length, and real phone voice notes (the quiet clips are simulated).
