Text-Based Video Editor
Edit a video or podcast by editing its transcript. Delete words to cut them, then remove filler words and long pauses in one click. Runs entirely in your browser, no upload, no signup.
The first run downloads the speech model (~200 MB, or ~77 MB without the graphics card), plus a word aligner (~67-90 MB) for English when your browser can use the graphics card. They stay in your browser, so the next file starts right away. Filler detection works best in English.
Stays on your device. Whisper transcribes in your browser and the cut is rendered here, so your video is never uploaded. Free, no signup, no watermark.
Get the next tool first
New local-AI tools and updates by email, plus a subscriber discount on the Mac app at launch. Unsubscribe anytime.
Edit a video by editing its transcript, without uploading it
Drop in a podcast episode, interview, talking-head video, or voice recording and get back a transcript you can edit like a document. Delete a sentence and it is gone from the video. Clear out the ums, uhs, and dead air in one click. Then export an MP4, an audio file, or subtitles. Everything runs in your browser. The file is never uploaded, there is no account, and there is no watermark.
Cloud editors make you upload the whole recording before you can cut a single word, and for a long episode that upload is often the slowest part of the job. This tool skips it.
How to edit a video by editing text
- Drop a video or audio file onto the tool and pick the spoken language. Talking content works best: podcasts, interviews, lectures, and voiceovers.
- Whisper transcribes it on your device, with a time for every word. For English, when your browser can use the graphics card, a second model then lines each word up with the audio, so your cuts land between words. The first run downloads the models; your browser caches them, so every later file skips the wait.
- Select words and press Delete to cut them. Cut words stay visible and struck through, so you can restore them, and undo and redo work the way you expect. Play the preview with cuts on to hear the edit.
- Press Remove fillers to cut every filler word and hesitation, and Shorten pauses to trim every long silence. Click a single pause to cut or keep just that one, or select any word to cut it by hand.
- Export the edit as MP4, M4A, MP3, or WAV, or take the subtitles (SRT, VTT) and the transcript (TXT) on their own.
What it finds for you
Filler words. Um, uh, erm, er, ah, hmm, and mm are marked in the transcript wherever Whisper wrote them down. For English, the tool also nudges Whisper to keep fillers in the text instead of quietly dropping them.
Hesitations the transcript missed. Speech models still skip plenty of fillers. A voice activity detector (Silero VAD) checks the audio for stretches of sound that no transcribed word covers, and marks each one as a hesitation you can cut like a word.
Long pauses. Any silence of 0.8 seconds or more between words shows up as a pause. Shorten pauses trims each one to about 0.3 seconds instead of removing it, so the speech still breathes.
Private by architecture, not by policy
There is no server doing the work, so there is nothing that could see your recording. That matters for unreleased episodes, client calls, internal meetings, and anything under an NDA. The models download once and your browser caches them: Whisper Base is about 200 MB on WebGPU, or about 77 MB on browsers that run it without the graphics card, the Silero voice detector is about 1.2 MB, and the wav2vec2 word aligner, used for English on browsers that run Whisper on the graphics card, is about 67 to 90 MB depending on the card. Whisper is the same speech model our audio transcription tool and AI video clipper use, so if you have run either in this browser, it is already there. It is also why the tool is free: we have no inference bill, so you get no meter.
Export formats and long recordings
MP4 export re-encodes the video at its original resolution and frame rate, with H.264 video and AAC audio, and you can burn captions in as you go. Every cut gets a tiny audio fade so the seams do not click. Audio exports come as M4A, MP3, or WAV.
In Chrome and Edge on a computer, you choose where to save before the export starts and the file streams straight to your disk, so even a multi-hour recording works. Other browsers, including Safari, Firefox, and every iPhone browser, hold the whole file in memory, so exports estimated over about 750 MB on a computer, or about 300 MB on a phone or tablet, are refused before they start. For those, switch to Chrome or Edge on a computer, or cut the recording shorter.
Honest limits
- Filler detection works best in English. Other languages transcribe fine and cutting works the same, but fewer fillers get caught.
- There are no speaker labels yet, so a two-person interview reads as one stream of paragraphs.
- Word times are close but not exact, and can be off by a tenth of a second or so. Play each cut back before you export, especially tight cuts between words.
- Your edits live in this tab only. Export before you close it; the page warns you if you try to leave with unsaved edits.
How it compares with Descript
Descript made transcript-based editing popular, and it is a far bigger product: multitrack timelines, screen recording, speaker detection, AI voice correction with Overdub, and team collaboration. It is a cloud service with a limited free plan and paid tiers, and your media is uploaded to its servers for transcription.
This tool does one part of that job, cutting a single recording by its transcript and cleaning out fillers and pauses, and it does it on your own device, free and without an account. If that is the part you need, you do not have to upload anything to get it.
Frequently asked questions
- Is it a free Descript alternative?
- For the core idea, yes: you edit a video or podcast by editing its transcript, and you can strip out filler words and long pauses in one click. It is free and unlimited with no signup, because the AI runs on your own computer, so there is no per-minute bill to pass on. Descript does much more, including multitrack editing, speaker labels, AI voice correction, and team collaboration. If you need a quick, private cleanup pass on a single recording, this covers it.
- Does my video get uploaded?
- No. The file never leaves your device. Transcription, filler detection, cutting, and export all happen locally in your browser. You can open your network tab and watch: after the one-time model download, nothing is sent anywhere.
- Can it remove um and uh?
- Yes. Filler words (um, uh, erm, er, ah, hmm, mm) are marked in the transcript, and Remove fillers cuts them all at once. Speech models often leave fillers out of the text entirely, so a second model, Silero VAD, listens for speech that no transcribed word covers and marks it as a hesitation you can cut like a word. Words like "like", "so", and "you know" are never touched, because they are often real words. Play the result back before you export: detection works best in English and can miss a filler or flag a short real word.
- What happens when I fix a word in the transcript?
- Deleting a word cuts that stretch of audio and video. Editing a word's spelling only changes the text, which is useful for names and terms the speech model misheard: the fix carries into the subtitle and transcript exports and the burned-in captions, but the audio stays as it was spoken.
- Which formats can I export?
- Video as MP4 (H.264 video with AAC audio), audio as M4A, MP3, or WAV, subtitles as SRT or VTT, and the edited transcript as plain text. Captions can be burned into the MP4 if you want them. Audio-only files, like a podcast episode, export to the audio, subtitle, and text formats.
- How long can my video be?
- There is no fixed limit on transcription. Speed depends on your graphics card: a recent fast one transcribes an hour of audio in about ten minutes, and a typical laptop takes longer. For export, Chrome and Edge on a computer write the file straight to your disk, so length does not matter. Other browsers build the file in memory first, so exports estimated over about 750 MB on a computer, or about 300 MB on a phone or tablet, are refused up front; use Chrome or Edge on a computer for those, or cut the recording shorter.
- Which languages does it support?
- Pick the spoken language before you start; English is the default. Whisper transcribes all the listed languages, and deleting words to cut them works the same in each. Filler and hesitation detection is tuned for English, so in other languages it catches fewer fillers.
- Does it work on iPhone?
- It works, but slowly. Phones run the speech model on the processor instead of the graphics card, and iOS limits how much memory a browser tab can use, so long recordings can run out of memory. iPhone browsers also cannot save straight to disk, so the in-memory export limit applies. For anything longer than a few minutes, use a computer.
- Are the models open source?
- Yes. Whisper Base (OpenAI) and Silero VAD are both MIT licensed. The word aligner used for English when your browser can use the graphics card, wav2vec2 Base 960h (Meta), is Apache 2.0 licensed. The cutting and encoding uses mediabunny, an open source WebCodecs toolkit, and MP3 export uses the LAME encoder.