Captions & transcription

The Captions tab in the left rail has three sub-tabs: Transcripts, Presets, and Settings. Transcription runs on your machine unless you pick the OpenAI engine, and the transcript it produces is what the agent uses to cut, search, and build story edits.

Generate captions

  1. Put a clip with speech on the timeline.
  2. Open Captions → Settings. Choose a Model and a Language (Auto detect, or English, Spanish, Italian, French, German, Portuguese, Russian, Japanese, Chinese).
  3. Click Generate captions. The first run of a local engine downloads its model; you will see “Downloading transcription model (141 MB)” for Fast. Progress then goes through Preparing model, Transcribing, Checking transcript, Generating captions.

Already have a subtitle file? Import takes .srt and .ass. Generate again reruns with the current settings.

Tip: If the button is disabled with “No audio detected”, the timeline has no clip with an audio track.

Engines

Transcription engines
ModelWhat it isNeeds
Fast (local)whisper.cpp, base model, with alignment and timing repair. Quick captions.Nothing. Bundled; 141 MB model on first use.
Best (local)Parakeet. Word-for-word, keeps retakes and fillers, native word timing. The default and the best choice for cutting.Python 3.10+ with onnxruntime; about 2.6 GB of model files on first use. Setup per OS on the prerequisites page.
Whisper (local)faster-whisper large-v3-turbo. Strong on accents and noisy audio; may merge repeated takes. Meant for clips up to about 10 minutes.Bundled on Windows. On macOS, Python 3.10+ with faster-whisper.
OpenAI (cloud)OpenAI's Whisper API. Audio is uploaded to OpenAI.An OpenAI key under Settings → API Keys. The option is disabled until one is saved.

Install steps for the Python engines are under Prerequisites → Transcription engines.

Grouping, gaps, position

  • Caption grouping: Fixed word count (with a Number of words field), Phrases, or Sentences.
  • Shorten word gaps: tightens pauses between words without touching speech. Set “When gap is longer than” and “Shorten to” in seconds, preview how many gaps match, then Apply.
  • Position: X, Y, and scale for the whole caption layer. If you moved captions by hand in the Transform tab, this panel says so and you reset there.
  • Depth for this caption: Behind person or In front of person, once the clip has a Cinematic depth matte.

Templates

Generate captions first, then click a card under Presets to apply its look to every caption.

Caption templates
TemplateStylePlan
SimpleDefaultFree
TikTok PillClassicFree
Word by WordSing-alongFree
TypewriterNarrationFree
CinematicElegantFree
DevinSignaturePlus and above
TrendyTrendingPlus and above
Neon HighlightTrendingPlus and above
Motion TitleSignaturePlus and above

Locked cards show a Plus badge; clicking one opens the plans view instead of applying. A template you have customised shows a Custom badge.

The template editor

The pencil on a card opens its editor. Fields, grouped:

  • Text: Size, Text Color, Stroke Color and Width, Line Height, Letter Spacing, Word Spacing, Highlight Color, Inactive Color, Payoff Color and presets, Secondary Color.
  • Layout: Casing (Default, None, Uppercase, Lowercase, Capitalize) and Width, the wrap width. Uppercase is how you get ALL CAPS; it never rewrites the text.
  • Background: colour, width, height, and a background animation.
  • Word Animation: None, Fade, Fade Up, Fade Down, Fade Left, Fade Right, Scale Bounce, Settle.
  • Shadow: on/off, colour, Blur, Offset X, Offset Y.
  • Trendy and Motion Title add Font Style, Emphasis Word, and an option to emphasise the longest word automatically.

Editing the transcript

The Transcripts sub-tab shows the words with their timing. It is also an editor for the video:

  • Select words and click Delete to remove that speech from the video. Deleted words stay in the list, grayed out and struck through; hover one and click Restore to put the footage back. Restores work from the most recent deletion backwards.
  • When the agent cuts, the same view becomes a cut review: grayed words are what will be removed. Click a word to keep it, double-click to seek.
  • Double-click a word to correct its text.

Cuts made here are the same timeline edits the agent makes, so a whole batch can be reverted from the chat panel.

Asking the caption agent

In chat, caption requests go to a dedicated caption agent that creates, regroups, restyles, and fixes captions and then checks its own work. Things it handles well:

  • “Add captions” or “Add captions in the Devin style”.
  • “Make the captions all caps” (sets casing, not a rewrite).
  • “Three words per caption” or “group by sentence”.
  • “Move the captions up a bit and make them smaller”.
  • “Make ‘revenue’ yellow in that line” or “add a fire emoji after ‘launch’”.
  • “Fix the spelling of my name in the captions”.

Caveat: Changing the words-per-caption or grouping rebuilds the caption track from the transcript and discards per-caption tweaks, per-word colours, and text edits. Do regrouping first, styling second.