Benchmarks

How fast Scenaristo transcribes, measured on real episodes.

These are the numbers behind the speed claims on this site. They were measured on four full length podcast episodes that are public on YouTube, so you can download the same files and run the same test on your own machine. We publish the machines and the method here, and a SHA-256 for every file in the results thread.

We chose real episodes rather than clean test audio on purpose. Two were recorded in a studio, two in front of a live audience with laughter, applause and people talking over each other. Processing time varies with the number of speakers and the audio conditions, and the table shows that variation instead of hiding it.

Results

EpisodeLengthSpeakersAudioLanguageMacBook Air (M1), 16 GBWindows 10, RTX A5000 *
Smosh Reads Reddit Stories"The Most Unpredictable Reddit Stories", watch on YouTube76 min 59 s5studioEnglish1 min 19 s2 min 43 s
Smosh Reads Reddit Stories"Can You Guess The Plot Twist?", watch on YouTube77 min 14 s5studioEnglish1 min 24 s2 min 43 s
La Ruinaepisode 205 (con Raúl Cimas), watch on YouTube73 min 19 s3live audienceSpanish2 min 06 s3 min 21 s
La Ruinaepisode 222 (con Andreu Buenafuente y Berto Romero), watch on YouTube95 min 02 s4live audienceSpanish2 min 46 s4 min 22 s

The Windows machine we measured carries an RTX A5000. The transcription decode runs as a single stream and uses only a fraction of the card: peak GPU utilization during these runs was about 50%. That is why a bigger GPU does not change the number, and why any NVIDIA Ampere GPU, the A5000 or a consumer card like the RTX 3070, lands in the same range.

Each time is measured from the moment you click Transcribe to the moment the full transcript is editable. It includes decoding the video, transcription, and placing every word on the timeline. It does not include speaker identification, which runs as a separate pass. Quoting the transcription step on its own would make us look faster than the experience actually is. Loading the model's weights is inside the time; downloading a model you do not have yet is not, so a first run on a cold cache is not reported as a slow one.

What that means in practice

An episode, transcribed before the coffee is cold.

1 min 19 s is roughly 58 minutes of audio processed for every minute you wait on the M1. The live audience episodes come in around 35 minutes of audio per minute of waiting. Scenaristo's own Benchmark card calls this RTFx and would print those two as 58× and 35×.

On either file you are reading an editable transcript inside three minutes. There is no upload, no queue, and no credit to spend.

58 minof audio per minute of waiting, studio episodes on the M1.
35 minof audio per minute of waiting, live audience episodes on the M1.
Test setup

Every condition, written down.

  • Mac. MacBook Air (M1), 16 GB, on battery.
  • Windows. Windows 10, NVIDIA RTX A5000 24 GB, 8 CPU cores.
  • Source files. 1920x1080 H.264 with AAC stereo audio, downloaded from YouTube at the highest available H.264 quality. Nothing was re-encoded, trimmed or cleaned before the test. Scenaristo never transcodes your footage without asking, and it did not here.
  • Network. Off for the whole corpus. Recorded as a condition of the runs rather than as a step: transcription is on-device, so the network is not in the loop either way.
Run it yourself

Check the numbers on your own machine.

Nothing here is a private test corpus. The four episodes are linked from the table above, and the results thread carries the rest: a SHA-256 for each file so you can confirm you are timing the same bytes we did, and how to make Scenaristo report its own timings rather than hold a stopwatch.

Two machines is not a lot of machines. If your numbers land far from ours on comparable hardware, that is worth knowing, and we would rather fix the table than defend it.

Results posted there stay in the thread. The table above is only ever what we measured ourselves, on machines we can re-run on demand. If you would rather not post publicly, then hello@scenaristo.com reaches us instead.

The budget, not the boast

A number we hold ourselves to.

On the baseline machines above, one hour of footage becomes an editable transcript in under 3 minutes, and usually well under. The slowest result in the table, a 95 minute episode recorded in front of a live audience, works out at 1 min 45 s per hour of footage on the Mac and 2 min 45 s per hour on the Windows machine. A release that cannot hold the budget on these files and these machines does not ship.

What we do not claim

The limits of this page.

We do not claim to be faster than any other product. We have not run their software on this hardware with these files, so we have no basis to say so. We publish our own numbers and our own method, and you can check them.

These figures are transcription times, not accuracy figures. Accuracy depends on the audio you give it and the language you speak, and we will not put a single percentage on it. Supported languages are listed on the models page.

Questions

The ones people ask about speed.

Does it need a GPU?
No. The M1 in the table above has no discrete GPU. A dedicated GPU makes it faster, as the Windows column shows, but it is not required. The minimum is 8 GB of memory, 16 GB recommended, on macOS 14 Sonoma or later on Apple silicon, or Windows 10 x64.
Does it get slower on long files?
Processing time grows with the length of the file, roughly in proportion. A 3 hour recording takes about twice as long as a 95 minute one. It does not get slower per minute, and there is no cap on length beyond your disk and memory.
Is anything sent to a server?
No. Transcription runs on your machine, so the network is not in the loop and nothing is uploaded. We ran the whole corpus with it off. Nothing you run on your machine is metered.
Why do the Spanish episodes take longer?
They were recorded in front of a live audience with more speakers and more overlap than the studio episodes. Audio conditions change processing time more than language does, and the table shows that variation rather than hiding it.

Last measured August 26, 2026. Results are re-measured on every release that touches the transcription pipeline, and a slower result blocks the release.

Nothing to queue for

Time it on your own footage
on your own machine.

macOS · Windows · no account, no upload