Features
What Rooster does, and what it does not do yet
In active development · Not yet released
Rooster records, transcribes, cuts, captions and exports video on your own Mac. Every section below opens with its status, because a feature list written before a product ships is worth exactly as much as its least honest line. Working today means working in a development build; in development means specified and not finished.
Recording suite
Working today
Rooster records camera, screen, or screen with camera straight into a project. The recorder holds the decisions a take actually needs and no more: which camera, which microphone, which display, the resolution and the canvas aspect ratio, framing guides, a mirror toggle and a countdown. Controls lock while capture is running and the elapsed time is on screen throughout.
A teleprompter with presentation controls sits beside the record button, so a scripted piece does not need a second device propped against the monitor.
Recordings are written to Downloads by default and an existing file is never silently overwritten. Nothing is uploaded to begin a recording, and nothing needs to be online for one to finish.
Local transcription
Working today
Transcription runs inside Rooster's own process, not as a call to a service. The default engine is Parakeet TDT v3, which on Rooster's own fixtures matched whisper.cpp large-v3-turbo for accuracy at 13 to 28 times the speed and a fraction of its peak memory, and which covers 25 European languages. That figure is ours, measured on those fixtures, and it belongs to that comparison rather than to any machine you happen to own.
whisper.cpp stays selectable for the roughly 99 languages Parakeet does not cover, with word-level timestamps, progress and cancellation. There is also an OpenAI-compatible bring-your-own-key option for anyone who would rather use a service they already pay for. That is the one transcription path where audio leaves the Mac, and it happens only if you deliberately configure it.
New recordings and imports transcribe automatically as the project is created, with a live transcribing state in the transcript panel and a failure that stays on screen with a retry rather than vanishing. Surfacing fine-grained progress in the editor is still on the list; today it shows an indeterminate state.
Transcript editing
Working today, with more in development
Click a word and the playhead moves to it. Select a run of words and the header shows the word count and the measured duration of the selection in timeline time, so it counts what an export would keep rather than what the transcript says. Words whose media has already been cut read muted and seek to the join.
Press Delete and the selection becomes one undoable ripple edit: the cut snaps to word boundaries and measured silence, and the gap closes across every track and every caption cue at once, so an overlay does not drift away from the words it was placed over. Correct is a separate action that fixes the text without touching the media, and it is deliberately separate, because inferring whether you meant to fix a word or cut it is the one mistake this surface cannot make. Corrections are journalled, so they survive a re-transcription, and the ones a new transcript cannot support are listed rather than quietly dropped.
In development: rendering measured silences as selectable gap tokens so a pause can be cut exactly like a word (the measurement itself is built, the tokens are not), cut markers with restore-removed-media, filler-word review, and re-recording a selected range with the built-in recorder.
Captions
Working today, with more in development
Four caption presets, previewed on the canvas the moment you pick one, with caption-safe guides so nothing lands under a platform's interface. Captions burn into the exported video through the same composition path the preview uses, which is why what you see and what renders are the same thing rather than two implementations that agree most of the time.
SRT and VTT sidecars are written from the same cues, with the line and card limits the presets use. Generating and removing captions are recorded as undoable edits rather than writes that discard your history.
In development: applying a style to one clip rather than the whole project, font, weight, size, text and active-word colour, background box, outline, shadow, alignment and vertical position, and per-card line limits. Speaker labels are a half-built case worth naming: the model carries them and the subtitle export writes them, but nothing yet decides who is speaking.
Studio Sound
In development
DeepFilterNet speech cleanup runs on the Mac, with a strength control and A/B playback so you can hear exactly what it changed before keeping it. It works today in development builds.
It is listed as unfinished for a packaging reason rather than a quality one: the bundled helper it uses still needs to be signed as part of the application, and the option to point it at an executable of your own has to come out, before Rooster could ship it through the App Store. Naming that here is more useful than a tick that quietly means something narrower.
Eye Contact
Working today, in beta
Eye Contact adjusts where you appear to be looking when you were reading a script rather than the lens. It is built on Apple's Vision pupil tracking, runs on your Mac, and has A/B playback so the correction is something you accept rather than something applied behind your back.
It is labelled beta in the app and on this page for the same reason: it is a per-frame estimate of where a person is looking, and on some footage it is better to leave it off.
Green Screen
Working today
Background blur, studio backgrounds, chroma green and your own images, with edge controls, all rendered on the Mac. No physical green screen is required for the segmentation modes.
Rendering effects locally costs time rather than credits, and the trade is worth stating plainly: on an M-series Mac, a five-second render of 1920x1080 at 60 fps took about 14 seconds, so a minute of video is roughly three minutes of rendering. A quality choice for the segmentation step, which is where nearly all of that time goes, is on the list.
Export
Working today
H.264 MP4 or HEVC MOV, at 720p, 1080p or the project's own size, with audio-only, SRT, VTT, plain transcript and Markdown available alongside the video. Progress and cancellation throughout, sleep prevented while a render runs, a failed export cleaning up its partial file, and Downloads as the default destination.
The render happens on your Mac. Saving a file does not publish anything, does not create a share link and does not require an account, and preview and export consume the same edit graph, so the file is the thing you were watching.
In development: preserving colour space, orientation and audio to video synchronisation across every source is an open item, as are the trim, move and layered tracks the timeline still needs.
B-roll and layered media
In development
This one is further along underneath than it looks from the outside. The edit model already has an overlay track kind, every clip already carries a full picture-in-picture transform, and the composition builder already composites an overlay above the primary video in both preview and export.
What is missing is the way in: adding media to an existing project, a bookmark for each asset so a second file reopens after a relaunch, importing images and audio rather than only video, a second visible timeline lane, drag and drop, and covering a transcript selection with the chosen clip. None of that is built, so B-roll is not something you can do in Rooster today.
Generated B-roll is a separate, later decision, recorded as a bring-your-own-key flow: your key, your chosen provider, no Rooster meter, and prompts leaving the Mac only to the provider you picked. It is specified and not started.
AI assistance
In development
The planned set is small and deliberately unglamorous: removing filler words with a review list before anything is cut, detecting and shortening long gaps, spotting likely retakes and offering the candidates for approval, suggesting B-roll search phrases from the transcript, and finding highlight ranges to spin into short-form cuts. None of it is built.
One rule shapes all of it. The model selects word or sentence spans from the timestamped transcript and never emits a timecode; the application owns every piece of timecode arithmetic and snaps each cut to word boundaries and measured silence. A language model that is allowed to invent timings will eventually invent one, and a cut in the wrong place is worse than no feature.
The second rule is review before apply. Nothing in this list modifies your timeline on its own, and everything it does is one undo away.
Things Rooster is not
Deliberately out of scope
Rooster is not everything a cloud editor does, and the shortest way to be useful is to say which parts. There is no collaboration, no comments and no cloud projects. There are no AI voices, no speech synthesis, no dubbing and no translation. There is no multicamera editing, no template or trend library, no stock catalogue and no publishing to a social platform. It is a Mac application on Apple silicon, and there is no Windows, iPad or browser version planned.
What is left is the loop a solo creator actually repeats: record, transcribe, cut by reading, caption, export. The solutions pages say what that means for four kinds of work, and each one lists the same exclusions in the terms of that job.
Where the work happens
Transcription, Studio Sound, Eye Contact, Green Screen, preview and export all run on your Mac. Your footage never leaves it unless you deliberately connect a cloud provider, and the qualifier is there because two such connections exist by design: downloading a transcription or language model, and supplying your own API key for an OpenAI-compatible transcription service. Both are listed on how it works and in the privacy policy.
That is also why nothing here is metered. Running the work on your machine means your usage costs us nothing, so there are no media minutes and no AI credits anywhere in the design, and no version of this page in which a longer recording costs more.