Editing, rebuilt around the words.
For spoken-word video, the transcript already is the edit decision list. Scenaristo treats it that way, and keeps a real timeline within reach for the moments that need a frame, not a sentence.
The transcript is the editor.
Every word carries its own timestamp, so the document and the footage are two views of one edit. Select a phrase, press delete, and the media ripples closed. Cuts land on word boundaries with a short crossfade so there's no click.
- Click a word to jump the playhead there instantly.
- One undo history spans text and timeline edits together.
- Read it like an essay, or switch to a script-style page layout.
- Search the transcript to jump straight to any phrase.
- Start editing before it finishes. Transcription streams in the background, so a long recording is editable from the first transcribed minute.
Editorial reader · 9,610 words · word-level timingA real timeline, for the frame-level moments.
Most cuts happen in text. But when one lands mid-breath, the classic timeline is right there: trim a boundary by the frame, blade a non-speech region, nudge a cut point. Whatever you do, the transcript stays perfectly in sync.
- Video, audio & text lanes, aligned to the same playhead.
- Waveform & zoom for precise, frame-accurate adjustments.
- Timeline-only cuts in silence show up as markers in the text.
Timeline view · synced transcript railQuiet features that do real work.
Strikethrough soft-delete
Preview a cut before committing. Struck text stays visible but grayed; playback and export skip it; one click restores it, losslessly.
Non-destructiveFiller-word cleanup
Every "um," "uh," "like," and "you know" is flagged. Review them one by one, or strike the entire set in a single click.
Bulk or per-instanceCaptions for free
Export SRT and VTT from the corrected, post-cut transcript, with cues grouped at natural punctuation, or burn them straight into the picture for Shorts. Your subtitles match the final edit exactly, no second pass.
SRT · VTTCorrection mode
Fix a mistranscribed name without touching a frame of video. A distinct mode keeps text corrections and cuts cleanly apart.
Text-onlyMulti-speaker labels
Interviews and panels split into renamable speaker turns, each voice re-identified consistently across the whole file, so jumping around takes seconds, not scrubbing.
On-device diarizationLoudness normalization
Master audio to a target (Podcast −16 LUFS, custom, or the −14 LUFS social target applied automatically to video exports), within ±0.5 LU, with a true-peak ceiling and the measured result reported back.
EBU R128 · LUFSChapters
Pin chapters on words in the transcript. Export a YouTube 00:00 list to paste, plus ID3 CHAP frames so podcast apps show them natively. Chapters move with your cuts.
Send to your NLE
Export an FCPXML 1.11 timeline, cuts, crossfades, and chapter markers preserved, to finish in Final Cut Pro, Premiere, or DaVinci Resolve. No media is re-encoded.
FCPXML handoffExport All, in one pass
Render MP4 at source resolution, encoded to the platforms' recommended H.264 spec, or MP3/WAV for podcasts, or bundle video, captions, and chapters behind a single dialog with a shared name. Duration equals source minus your cuts.
MP4 · MP3 · WAVRemove silences
Over-long pauses are detected alongside filler words. Strike them all in one click, or step through each gap in the reader.
Silence detectionClips & vertical export
Select a range, press ⌘⇧K, and export it as its own 9:16 or 1:1 MP4 with a re-timed SRT: crop overlay, platform safe-area guides, and on-device auto-framing included.
9:16 · 1:1 · ⌘⇧KBurned-in captions
Composite captions into the frames for muted autoplay, in a fixed high-legibility style identical to the sidecar SRT, previewed exactly as they ship.
Open captionsDeliverable Modes
On first import, say what you're making (Podcast, Long-form, or Short-form) and the right export presets and workspace layout are set in one step. Always changeable, never re-imposed.
Podcast · Long-form · Short-formOne project, one file
A project is a single .scenar file. Back it up, duplicate it, move it like any document. Every edit is a non-destructive layer inside it, never a change to your source media.
On-device, and honest about it.
Transcription runs on best-in-class open speech models, downloaded on demand and checksum-pinned. Nothing ships bundled in the installer. Pick the fast default, or reach for a heavier model when accuracy calls for it. See the full model family.
Import
MP4, MOV, MKV and WebM (VP9/Opus) for video. MP3, WAV, AAC, and M4A for audio. Drag-and-drop; the clip lands on the timeline before transcription even starts.
Transcription models
Nothing is bundled. The default Parakeet model and every alternative download on demand with pinned checksums, so the app stays small and you only fetch what you use.
macOS
macOS 11 Big Sur or later on Apple silicon. 8 GB of memory minimum, 16 GB recommended. Encode and decode ride the OS media engines, so exports are fast without a render farm. Projects live in a library with Recents and Save As, and transcription runs in the background with live progress on long files.
Windows
The same editor on Windows 10 x64, with a code signed installer and portable project files that open on macOS too. 8 GB of memory minimum, 16 GB recommended; discrete GPU recommended, not required. Download for Windows.
See your edit before you
touch a timeline.
Free today · no account required