How fast Scenaristo transcribes, measured on real episodes.
These are the numbers behind the speed claims on this site. They were measured on four full length podcast episodes that are public on YouTube, so you can download the same files and run the same test on your own machine. We publish the machines and the method here, and a SHA-256 for every file in the results thread.
We chose real episodes rather than clean test audio on purpose. Two were recorded in a studio, two in front of a live audience with laughter, applause and people talking over each other. Processing time varies with the number of speakers and the audio conditions, and the table shows that variation instead of hiding it.
Results
| Episode | Length | Speakers | Audio | Language | MacBook Air (M1), 16 GB | Windows 10, RTX A5000 * |
|---|---|---|---|---|---|---|
| Smosh Reads Reddit Stories"The Most Unpredictable Reddit Stories", watch on YouTube | 76 min 59 s | 5 | studio | English | 1 min 19 s | 2 min 43 s |
| Smosh Reads Reddit Stories"Can You Guess The Plot Twist?", watch on YouTube | 77 min 14 s | 5 | studio | English | 1 min 24 s | 2 min 43 s |
| La Ruinaepisode 205 (con Raúl Cimas), watch on YouTube | 73 min 19 s | 3 | live audience | Spanish | 2 min 06 s | 3 min 21 s |
| La Ruinaepisode 222 (con Andreu Buenafuente y Berto Romero), watch on YouTube | 95 min 02 s | 4 | live audience | Spanish | 2 min 46 s | 4 min 22 s |
The Windows machine we measured carries an RTX A5000. The transcription decode runs as a single stream and uses only a fraction of the card: peak GPU utilization during these runs was about 50%. That is why a bigger GPU does not change the number, and why any NVIDIA Ampere GPU, the A5000 or a consumer card like the RTX 3070, lands in the same range.
Each time is measured from the moment you click Transcribe to the moment the full transcript is editable. It includes decoding the video, transcription, and placing every word on the timeline. It does not include speaker identification, which runs as a separate pass. Quoting the transcription step on its own would make us look faster than the experience actually is. Loading the model's weights is inside the time; downloading a model you do not have yet is not, so a first run on a cold cache is not reported as a slow one.
An episode, transcribed before the coffee is cold.
1 min 19 s is roughly 58 minutes of audio processed for every minute you wait on the M1. The live audience episodes come in around 35 minutes of audio per minute of waiting. Scenaristo's own Benchmark card calls this RTFx and would print those two as 58× and 35×.
On either file you are reading an editable transcript inside three minutes. There is no upload, no queue, and no credit to spend.
Every condition, written down.
- Mac. MacBook Air (M1), 16 GB, on battery.
- Windows. Windows 10, NVIDIA RTX A5000 24 GB, 8 CPU cores.
- Source files. 1920x1080 H.264 with AAC stereo audio, downloaded from YouTube at the highest available H.264 quality. Nothing was re-encoded, trimmed or cleaned before the test. Scenaristo never transcodes your footage without asking, and it did not here.
- Network. Off for the whole corpus. Recorded as a condition of the runs rather than as a step: transcription is on-device, so the network is not in the loop either way.
Check the numbers on your own machine.
Nothing here is a private test corpus. The four episodes are linked from the table above, and the results thread carries the rest: a SHA-256 for each file so you can confirm you are timing the same bytes we did, and how to make Scenaristo report its own timings rather than hold a stopwatch.
Two machines is not a lot of machines. If your numbers land far from ours on comparable hardware, that is worth knowing, and we would rather fix the table than defend it.
Results posted there stay in the thread. The table above is only ever what we measured ourselves, on machines we can re-run on demand. If you would rather not post publicly, then hello@scenaristo.com reaches us instead.
A number we hold ourselves to.
On the baseline machines above, one hour of footage becomes an editable transcript in under 3 minutes, and usually well under. The slowest result in the table, a 95 minute episode recorded in front of a live audience, works out at 1 min 45 s per hour of footage on the Mac and 2 min 45 s per hour on the Windows machine. A release that cannot hold the budget on these files and these machines does not ship.
The limits of this page.
We do not claim to be faster than any other product. We have not run their software on this hardware with these files, so we have no basis to say so. We publish our own numbers and our own method, and you can check them.
These figures are transcription times, not accuracy figures. Accuracy depends on the audio you give it and the language you speak, and we will not put a single percentage on it. Supported languages are listed on the models page.
The ones people ask about speed.
Does it need a GPU?
Does it get slower on long files?
Is anything sent to a server?
Why do the Spanish episodes take longer?
Last measured August 26, 2026. Results are re-measured on every release that touches the transcription pipeline, and a slower result blocks the release.
Time it on your own footage
on your own machine.
macOS · Windows · no account, no upload