What transcript-first video editing actually changes
Most video editing tools present the same object: a horizontal timeline with clips on it. You scrub, you find the edit point, you trim, you nudge, you check the audio did not click. The interface is a faithful model of what film editing was when it involved actual film, and for a lot of work it is still the right model.
It is a poor model for talking. If the video is a person explaining something to a camera, the structure of the piece is not visual at all. It is the order of the sentences. And a timeline gives you no way to see sentences.
The unit of work changes
Transcript-first editing swaps the primary view. Instead of a waveform you get the words, and cutting a sentence out of the text cuts it out of the video. The picture and the sound follow.
That sounds like a convenience feature. It is really a change in what the software considers the unit of work. On a timeline the unit is a frame range. In a transcript the unit is a phrase, which happens to be the unit you were already thinking in when you decided the second take was better than the first.
The gap between those two units is where most of the tedium of editing talking footage lives. You know what you want to remove. Translating that into a selection on a timeline is mechanical labour that contributes nothing.
Which decisions get cheap
The interesting effect is not that individual edits get faster. It is that certain decisions stop being expensive enough to avoid.
Reordering two paragraphs of an explanation is a genuine chore on a timeline: find both regions, cut, move, close the gap, fix the audio transitions, watch it back. So people do not do it. They accept the order they happened to say things in, because restructuring costs more than the improvement is worth.
When the transcript is the edit, reordering is a cut and paste. It costs about as much as moving a paragraph in a document. Suddenly it is worth doing, and the finished video is better in a way that has nothing to do with editing speed - it is better because a decision that used to be too expensive became free.
The same applies to cutting. If removing a rambling thirty seconds is a five-second operation, you remove it. If it is a two-minute operation, you talk yourself into keeping it.
Where the timeline is still right
This is not an argument that timelines are obsolete. A transcript is a terrible interface for anything where the structure is visual:
- Cutting to music, where the edit points are beats and not words
- Multi-camera work with framing as the reason for each cut
- Anything with significant motion graphics or compositing
- Silent footage, obviously, which has no transcript at all
A tool that only offers a transcript is making the opposite mistake to a tool that only offers a timeline. The useful arrangement is a transcript as the primary view for talking footage, with a real timeline underneath for the work that needs one.
The accuracy question
The obvious objection is that transcript editing is only as good as the transcript. If the words are wrong, the edit points are wrong.
That was a fair objection for a long time. Speech recognition was accurate enough to search but not accurate enough to cut against, particularly at word boundaries, which is exactly where a transcript editor needs precision. A model that gets the sentence right but the word timings approximately right will produce edits that clip the start of words.
Word-level timestamps, rather than sentence-level ones, are the thing that makes this work. It is a less glamorous capability than the transcription itself and it matters more.
What we are building
Rooster is a Mac editor built on this idea, with transcription running on your own machine rather than a server. It is in active development and not released yet, so this is a description of a design rather than a product tour.
The reason for writing it down now is that the design decision came first and everything else follows from it. If you are choosing tools for talking-head video, the question worth asking is not which has more features. It is which one matches the unit you are actually thinking in.