FeaturesSolutionsPricingBlog

Capability

Captions from your own transcript

In active development · Not yet released

Captions are built from the transcript's timings, so there is no second transcription and no upload. Four presets preview on the canvas as you pick one, then adjust: typography, colour, outline, shadow, placement, and how many characters a line holds. Export burns them into the video and writes SRT and VTT.

How captions are made

Captions come from the transcript, which means the video has to be transcribed first. The caption panel says exactly that rather than leaving a gallery of styles mysteriously inert.

There are four presets. Classic is a readable box caption and the safe default for any video. Clean Paragraph drops the background and takes a larger line budget, for talking-head video where the picture should stay visible. Bold Highlight is short, heavy cards for short-form social video watched without sound. Karaoke highlights each word as it is spoken, driven by the per-word timings the local engines produce.

Each one previews on the canvas the moment you pick it, and the View menu can draw safe-area guides so nothing lands where a platform's own interface will cover it. Those guides are drawn on the canvas only and never reach a render.

Captions burn into the exported video through the same composition path the preview uses, which is why what you see and what renders are the same thing rather than two implementations that agree most of the time. SRT and VTT sidecars are written from the same cues, with the same line and card limits. Generating captions and removing them are both recorded as undoable edits rather than writes that discard your history.

What is on screen

  • The preset gallery

    Four styles, each drawn as a thumbnail as well as named, with the current one ticked. Picking one applies it and rebuilds the cards. A style you have adjusted keeps its card selected and the panel says so, with the way back to the preset in one place.

  • Typography

    A font picker, a weight picker and a size slider, applied to every card in the project.

  • Colour

    The text colour, the active word's colour where a style highlights one, a Highlight each word toggle, and a background box you can switch off or recolour. Every colour well supports opacity.

  • Legibility

    An outline with its own width and colour, and a shadow with a radius, an offset and a colour. Zero means off for either. This is what keeps an unboxed caption readable over a busy picture.

  • Placement

    Left, centre or right alignment, and a vertical position slider running from higher in the frame to lower.

  • Cards

    Characters per line and lines per card, as steppers. Changing either reshapes the cues rather than restyling them, and the panel says so before you touch it.

  • Remove captions

    A link under the style controls, recorded as an undoable edit like generating them was.

Every capability Rooster ships is listed on the features page, and how the work is split between your Mac and everything else is on how it works.

What it does not do

  • Word-by-word highlighting needs word timing, which the local engines produce. Where a transcript has none the toggle is disabled and the panel explains why, rather than lighting up nothing.
  • Captions travel as one set. A change styles every card in the project; there is no per-card or per-word styling.
  • Nothing in Rooster decides who is speaking, so captions carry no speaker names of their own.
  • The panel lists the first 40 cues and says how many there are in total. It is a check on where the cards break, not a caption editor - the words themselves are corrected in the transcript.
  • Captions are burned into the exported video. If you need them switchable for the viewer, that is what the SRT or VTT sidecar is for.

This list is here because the alternative is a reader installing Rooster and finding out. If one of these is the job you need done, the alternatives page ranks tools that do it.

Get one email when Rooster ships

Rooster is in active development. Leave an address and we will send a single launch announcement - no newsletter, no drip sequence, nothing else.

Used for the launch announcement and nothing else.

Optional, and it genuinely helps: a search, a link, a person, an AI answer - whatever it was.

Frequently asked questions

Do captions need a second transcription or an upload?

Neither. They are built from the transcript already in the project, using its timings, on your Mac. If the video has not been transcribed yet, the caption panel says so and points you at doing that first.

Can I change the font, the colour and where captions sit?

Yes. Font, weight and size; the text colour, the active word's colour and a background box with its own colour; an outline and a shadow, each of which can be set to zero to turn it off; and alignment with a vertical position slider.

Does Rooster write SRT and VTT files?

Yes, from the same cues as the burned-in captions, with the same characters-per-line and lines-per-card limits. Both are export options alongside the video, audio-only, plain transcript and Markdown.

Why is "Highlight each word" greyed out?

Because that transcript has no word timing. Word-by-word highlighting needs it, so rather than draw a control that would do nothing, Rooster disables it and says which piece is missing.

Can I style one caption card differently from the rest?

No. Captions travel as one set in Rooster, so a style change applies to every card in the project.

The rest of the loop

  • Screen recording on your Mac

    A display or one window, with Mac audio and your microphone, cropped to the shape you are publishing in.

  • Transcription that runs on your Mac

    Parakeet or whisper.cpp running in Rooster's own process, with the word timings the rest of the editor is built on.

  • Removing filler words

    Every hesitation in the transcript, played in context, cut as one undoable step when you approve it.

  • Studio Sound

    DeepFilterNet 3 on your Mac, with a strength control and A/B playback before anything is kept.

  • Eye Contact

    A restrained, on-device gaze correction with three strengths and A/B playback. Beta, and labelled so in the app.

  • The teleprompter in Rooster's recorder

    A script window opened from the recorder, with mirror and flip for a beam splitter, and a choice about whether it appears in the take.

Local, transcript-first video editing for Mac.

In active development. Not released yet.

Product

  • Features
  • Solutions
  • Pricing
  • How it works
  • Support

Compare

  • All comparisons
  • Loom alternatives
  • Descript alternatives
  • Descript vs Riverside
  • Descript vs CapCut
  • Descript pricing

Resources

  • Blog
  • FAQ
  • Tools

Legal

  • Privacy
  • Terms
  • Acceptable use
  • Licences

Ask AI about Rooster

Opens your assistant with the question already written. Answers come from the model, not from us.

  • ChatGPT
  • Claude
  • Perplexity
  • Grok

© 2026 Rooster.

Transcription, editing and export run on your own Mac.