Features

Editing, rebuilt around the words.

For spoken-word video, the transcript already is the edit decision list. Scenaristo treats it that way, and keeps a real timeline within reach for the moments that need a frame, not a sentence.

Text-based editing

The transcript is the editor.

Every word carries its own timestamp, so the document and the footage are two views of one edit. Select a phrase, press delete, and the media ripples closed. Cuts land on word boundaries with a short crossfade so there's no click.

  • Click a word to jump the playhead there instantly.
  • One undo history spans text and timeline edits together.
  • Read it like an essay, or switch to a script-style page layout.
  • Search the transcript to jump straight to any phrase.
  • Start editing before it finishes. Transcription streams in the background, so a long recording is editable from the first transcribed minute.
A 92-minute recording laid out as an editable essay in the transcript reader, 9,610 words with word-level timing.Editorial reader · 9,610 words · word-level timing
Hybrid by design

A real timeline, for the frame-level moments.

Most cuts happen in text. But when one lands mid-breath, the classic timeline is right there: trim a boundary by the frame, blade a non-speech region, nudge a cut point. Whatever you do, the transcript stays perfectly in sync.

  • Video, audio & text lanes, aligned to the same playhead.
  • Waveform & zoom for precise, frame-accurate adjustments.
  • Timeline-only cuts in silence show up as markers in the text.
The timeline view with struck words mirrored as closed cuts across the video and audio lanes, transcript synced in the side rail.Timeline view · synced transcript rail
The rest of the toolkit

Quiet features that do real work.

Strikethrough soft-delete

Preview a cut before committing. Struck text stays visible but greyed; playback and export skip it; one click restores it, losslessly.

Non-destructive

Filler-word cleanup

Every "um", "uh", "like", and "you know" is flagged. Review them one by one, or strike the entire set in a single click.

Bulk or per-instance

Captions for free

Export SRT and VTT from the corrected, post-cut transcript, with cues grouped at natural punctuation, or burn them straight into the picture for Shorts. Your subtitles match the final edit exactly, no second pass.

SRT · VTT

Correction mode

Fix a mistranscribed name without touching a frame of video. A distinct mode keeps text corrections and cuts cleanly apart.

Text-only

Multi-speaker labels

Interviews and panels split into renamable speaker turns, each voice re-identified consistently across the whole file, so jumping around takes seconds, not scrubbing.

On-device diarization

Loudness normalisation

Master audio to a target (Podcast −16 LUFS, custom, or the −14 LUFS social target applied automatically to video exports), within ±0.5 LU, with a true-peak ceiling and the measured result reported back.

EBU R128 · LUFS

Chapters

Pin chapters on words in the transcript. Export a YouTube 00:00 list to paste, plus ID3 CHAP frames so podcast apps show them natively. Chapters move with your cuts.

YouTube · ID3 CHAP

Send to your NLE

Export an FCPXML 1.11 timeline, cuts, crossfades, and chapter markers preserved, to finish in Final Cut Pro, Premiere, or DaVinci Resolve. No media is re-encoded.

FCPXML handoff

Export All, in one pass

Render MP4 at source resolution, encoded to the platforms' recommended H.264 spec, or MP3/WAV for podcasts, or bundle video, captions, and chapters behind a single dialog with a shared name. Duration equals source minus your cuts.

MP4 · MP3 · WAV

Remove silences

Over-long pauses are detected alongside filler words. Strike them all in one click, or step through each gap in the reader.

Silence detection

Clips & vertical export

Select a range, press ⌘⇧K, and export it as its own 9:16 or 1:1 MP4 with a re-timed SRT: crop overlay, platform safe-area guides, and on-device auto-framing included.

9:16 · 1:1 · ⌘⇧K

Burned-in captions

Composite captions into the frames for muted autoplay, in a fixed high-legibility style identical to the sidecar SRT, previewed exactly as they ship.

Open captions

Deliverable Modes

On first import, say what you're making (Podcast, Long-form, or Short-form) and the right export presets and workspace layout are set in one step. Always changeable, never re-imposed.

Podcast · Long-form · Short-form

One project, one file

A project is a single .scenar file. Back it up, duplicate it, move it like any document. Every edit is a non-destructive layer inside it, never a change to your source media.

.scenar · non-destructive
Under the hood

On-device, and honest about it.

Transcription runs on best-in-class open speech models, downloaded on demand and checksum-pinned. Nothing ships bundled in the installer. Pick the fast default, or reach for a heavier model when accuracy calls for it. See the full model family.

Import

MP4, MOV, MKV and WebM (VP9/Opus) for video. MP3, WAV, AAC, and M4A for audio. Drag-and-drop; the clip lands on the timeline before transcription even starts.

Transcription models

Nothing is bundled. The default Parakeet model and every alternative download on demand with pinned checksums, so the app stays small and you only fetch what you use.

macOS

macOS 11 Big Sur or later on Apple silicon. 8 GB of memory minimum, 16 GB recommended. Encode and decode ride the OS media engines, so exports are fast without a render farm. Projects live in a library with Recents and Save As, and transcription runs in the background with live progress on long files.

Windows

The same editor on Windows 10 x64, with a code signed installer and portable project files that open on macOS too. 8 GB of memory minimum, 16 GB recommended; discrete GPU recommended, not required. Download for Windows.

Try it on your next recording

See your edit before you
touch a timeline.

Free today · no account required