The ASR model family

Five speech models, one transcript.

Text-based editing is only as good as the words underneath it. Scenaristo runs a small family of best-in-class automatic speech recognition models on your own machine, each picked for a low word error rate, fast turnaround, and the thing every cut depends on: accurate word-level timing.

What we select for

Accurate words, in time, on-device.

Every model in the family clears the same three bars before it ships in the app.

Low word error rate

Fewer mistranscriptions means fewer corrections, and a transcript you can trust enough to edit by deleting text.

WER-first

Fast turnaround

Transcription runs in the background and finishes quickly, so a long recording is editable in minutes, not coffee breaks.

Real-time class

Word-level timing

Each word carries its own timestamp. That is what lets a deleted phrase ripple the footage closed on a clean boundary.

Per-word stamps
Ranked by speed

Pick the trade-off your audio needs.

Fastest at the top, broadest coverage at the bottom. Scenaristo defaults to Parakeet and lets you reach for the others when accuracy or a rarer language calls for it. Every model downloads on demand, pinned by checksum, stored locally, never fetched twice.

#1Fastest

Parakeet TDT

Default modelNVIDIA
nvidia/parakeet-tdt-0.6b-v3

The default. A streaming-friendly transducer that turns an hour of audio around in a fraction of the time, with word-level timestamps baked in.

Token-and-Duration Transducer, built for real-time throughput.

25languages
#2Very fast

Canary

Accuracy pickNVIDIA
nvidia/canary-1b-v2

When accuracy matters more than raw speed. A heavier model that trades a little throughput for cleaner transcripts on tough audio.

A larger encoder for the lowest word error rate of the set.

25languages
#3Fast

Cohere Transcribe

Cohere Labs
CohereLabs/cohere-transcribe-03-2025

A well-rounded option focused on the most common languages, with reliable timing and punctuation out of the box.

Strong on clean speech across major business languages.

14languages
#4Broadest

Whisper Large v3 Turbo

Widest coverageOpenAI
openai/whisper-large-v3-turbo

The fallback for everything else. When your audio is in a language the faster models do not cover, Whisper almost certainly does, and Turbo gets you there quickly.

A distilled decoder, markedly faster than large-v3 for a small accuracy cost.

99languages
#5Thorough

Whisper Large v3

Accuracy rungOpenAI
openai/whisper-large-v3

The accuracy rung of the Whisper ladder. Slower than Turbo, but the strongest option on noisy audio and the long tail of languages.

The full, un-distilled model, reach for it on the hardest audio.

99languages

Model identifiers reference their public Hugging Face repositories. Parakeet, Canary, Cohere Transcribe, and Whisper are trademarks of NVIDIA, Cohere, and OpenAI respectively. Availability and exact versions may change as the family is updated.

How the family works together

A sensible default, with headroom when you need it.

Most recordings never need more than the default. Parakeet is fast, accurate, and covers the languages most spoken-word video is made in, so for the typical interview or talking-head clip, it just works.

When a file is noisy, multilingual, or in a language outside the faster models' range, you can switch models on download and re-run transcription. The edit, the cuts, and the timing all stay intact.

On-deviceEvery model runs locally. Audio never leaves your machine.
SwappableChange models per project and re-transcribe in place.
Up to 99languages covered across the full family.
Better words, better edits

State-of-the-art transcription,
running on your desk.

On-device · no account · no upload