---
title: "Make a Captioned Product Demo with Local Mac Voiceover"
description: "Turn a real Mac screen clip and local narration into a captioned product demo. Use explicit cues, preserve originals, and verify the finished MP4."
canonical: "https://www.murmurtts.com/blog/resources/captioned-product-demo-local-voiceover-mac"
---
[Murmur](https://www.murmurtts.com/)/[Resources](https://www.murmurtts.com/blog/resources)/Make a Captioned Product Demo with Local Mac Voiceover

Guide

# Make a Captioned Product Demo with Local Mac Voiceover

Turn a real Mac screen clip and local narration into a captioned product demo. Use explicit cues, preserve originals, and verify the finished MP4.

Murmur·Published October 9, 2026·8 min read

On this page

Record one action the narration can explainGenerate a local narration takeDraft captions, then review their actual cuesPrepare the files and render a fresh movieCheck the finished export and its revisionGive an assistant a bounded production briefFrequently asked questionsSources

[All resource guides](https://www.murmurtts.com/blog/resources)

Choose this workflow to make a product demo from a real Mac screen recording, local narration, and timed captions. Show one visible action, write a line that describes it, generate the WAV, review the caption cues, then assemble a new MP4. The useful result is a clip whose screen, speech, and captions agree about what happened.

The worked example opens Murmur’s model selector. Download the [editable workflow kit](https://www.murmurtts.com/downloads/captioned-product-demo-workflow.txt) and the [3.7-second rendered example](https://www.murmurtts.com/downloads/captioned-product-demo-example.mp4). The kit creates a new local source folder; it does not install packages or run commands for you. For a longer script, start with the [product-demo script template](https://www.murmurtts.com/blog/product-demo-script-template).

## Record one action the narration can explain

Apple's [screen-recording guide](https://support.apple.com/en-au/guide/quicktime-player/qtp97b08e666/mac) starts with QuickTime Player, File, New Screen Recording. Choose a capture area that keeps the control and its result readable. Record a deliberate action with a brief steady view before and after it. Save the original separately from any trimmed or cropped copy.

For this example, we selected seconds 6–11 of an existing Murmur product-tutorial recording. A 960 by 540 close-up from the upper-right portion of its 1920 by 1206 image became a 1280 by 720 clip. The selected range excludes the history sidebar. Its selector opens and stays visible; no model change or speech generation is shown. The recording’s exact app build is unknown, and visible model names describe that capture rather than today’s catalog.

> Open the model selector to see which speech engine is active.
>
> Original spoken narration for the worked demo

Keep the cropped control above the caption area. Review the entire selected range for unintended text and unrelated actions. If the screen continues into another operation before the line finishes, choose a better range or revise the line. The kit does not loop a short recording to fill the narration.

## Generate a local narration take

Save only the quoted sentence in narration.txt inside a new working directory. Confirm that the CLI is connected and the chosen model and voice are available. The [automation documentation](https://www.murmurtts.com/docs/automation) covers installation; the [MCP setup guide](https://www.murmurtts.com/blog/resources/connect-murmur-mcp-local-voiceover) covers assistant connections.

```
murmur status --json
murmur models --json
murmur voices --model qwen3-base --json
murmur generate --input narration.txt --model qwen3-base --voice Ryan \
  --language EN-US --speed 1.0 --output narration.wav --json
```

On October 9, 2026, one completed CLI job generated a 3.68-second, mono, 24 kHz PCM WAV using qwen3-base and Ryan through a development helper reporting Murmur 1.0.18. The output decoded successfully. Compatibility with a released commercial binary was not verified. Select an installed model and its own available voice on your Mac; the recording’s menu does not determine that choice.

Keep the completed WAV as the editing input. A file appearing on disk is not enough: confirm the job completed and the file contains a valid audio stream. “Unlimited local generation” still consumes hardware time, memory, and disk space. It does not include unlimited renderer capacity or licenses for third-party tools. Check [Remotion's license terms](https://www.remotion.dev/docs/license/pricing) for your use.

## Draft captions, then review their actual cues

We ran [whisper.cpp](https://github.com/ggml-org/whisper.cpp) 1.9.4 locally with base.en on a separate 16 kHz PCM copy. No script or prompt was supplied to the recognizer. Its normalized output matched all 11 script words. Machine timestamps suggested two phrases: “Open the model selector” from 0.19 to 1.33 seconds, then “to see which speech engine is active.” from 1.33 to 3.60. These are estimates for this take, not reusable timings for a new generation. The assembly kit itself does not transcribe audio.

Copy caption-review.example.json into a new captions.json. Replace the example duration and hash with measurements from narration.wav, then enter cues with startSeconds, endSeconds, and text. Use the exact spoken wording. The sample template's timings are placeholders, not timing evidence for this sentence.

```
shasum -a 256 narration.wav
ffprobe -v error -show_entries format=duration \
  -of default=noprint_wrappers=1:nokey=1 narration.wav
```

Review each cue against the audio and the visible action. W3C's [caption guidance](https://www.w3.org/WAI/media/av/captions/) calls for accurate speech and meaningful non-speech information; automatic text needs checking. Compare the beginning and end of every cue with the take. A matching transcript does not by itself establish correct timing or a comfortable reading pace.

Keep review.state pending until the named review has occurred. Our example records an automated transcript comparison plus visual engineering review, explicitly excluding human listening and guaranteed word alignment. The ordinary render requires approved and a reviewer name; this supplied record cannot independently establish review quality. Listen to your final take before using it in a promotional deliverable. Important visual information should also be understandable in narration, following [W3C’s description guidance](https://www.w3.org/WAI/media/av/description/); captions alone describe the audio.

## Prepare the files and render a fresh movie

Extract the download with Python 3 into a new folder, then read START_HERE.md. The working environment used Node 22.23.1, FFmpeg 9.0.2, Remotion 4.0.534 and React 19.3.0. Install the pinned packages and a Chrome Headless Shell separately. Set REMOTION_BROWSER_EXECUTABLE to its actual executable path. The renderer refuses a missing browser path; it does not download one during rendering. Initial dependencies and speech or transcription models can require network access.

```
python3 captioned-product-demo-workflow.txt demo-kit-01
cd demo-kit-01
npm install --registry=https://registry.npmjs.org --no-audit --no-fund
npx remotion browser ensure --chrome-mode=headless-shell
```

Keep the generated package-lock.json and use npm ci for later installs. Follow the browser command’s output to locate the executable, then set the environment variable described in START_HERE.md. Use a model weight file from the transcription project’s official download instructions if you choose local ASR; the kit does not bundle weights. The initial setup is separate from the local rendering stage.

Cut and crop the screen recording before preparation. The composition displays the supplied video at normal speed without further cropping, mutes its captured audio, and adds the local WAV. It uses [Remotion Video](https://www.remotion.dev/docs/media/video) and [Audio](https://www.remotion.dev/docs/media/audio); explicit cues are drawn by the composition. The [Remotion caption guide](https://www.remotion.dev/docs/captions/) covers other caption approaches.

```
node prepare.mjs screen.mp4 narration.wav captions.json run-01
node render.mjs run-01 captioned-demo.mp4
```

| Step | The kit checks | Your remaining review |
| --- | --- | --- |
| Prepare | Local files, audio hash/duration, ordered non-overlapping cues | Words and timing match the take |
| Fit | Screen clip covers the rounded narration timeline | The action remains visible long enough |
| Render | Prepared hashes, full decode, frame count, audio stream | Captions are readable and agree with the screen |
| Save | New run directory, movie, and sidecar names | Inspect any installed output before retrying |

Use new names for every revision. The prepared manifest binds the copied screen, narration, cues, and props by hash. If any prepared input changes, render stops. The target composition is 1280 by 720 at 30 fps with H.264 video, AAC audio, and captions burned into the picture. The existing [Remotion voiceover guide](https://www.murmurtts.com/blog/resources/remotion-local-voiceover-mac) explains the underlying audio-duration integration.

## Check the finished export and its revision

The finished example contains 111 frames at 30 fps, with a 3.700-second video track and a 3.712-second H.264/AAC container. The entire movie decoded successfully. We inspected the frames immediately before and after both caption changes: first cue on frames 6–39, second on 40–107, then no caption. The boxes remain below the menu. This is a short horizontal desktop example; reframe and retest type sizes for a vertical reel. Decode success and automated transcript agreement do not establish listening quality or accessibility conformance.

The kit rounds the narration timeline up to a frame; encoded audio or container duration can differ slightly. Recheck the final soundtrack and caption timing. Any new narration invalidates the old hash-bound cue record: transcribe, retime, review, and prepare again. Do not stretch old cues automatically. The [promo-reel guide](https://www.murmurtts.com/blog/resources/ai-promo-reel-local-voiceover-mac) is useful for title-card sequences; this workflow adds real product footage and timed captions.

## Give an assistant a bounded production brief

ChatGPT or Claude can help write narration and a shot plan. Executing local commands requires an assistant host with file and terminal access on your Mac. No assistant model is claimed to have executed this example. A useful script-only prompt is:

> Write one narration sentence for this observed screen action: [describe only the captured action]. Audience: [viewer]. Name the control and what the viewer can see after opening it. Return the sentence, its matching shot, and any unsupported claim to remove. Do not invent product behavior.
>
> Original script-planning prompt

> Use the local captioned-demo kit with these reviewed files: [screen], [narration], [cue record]. Verify hashes and durations, prepare a new run, render to a new output, and return the movie and report paths. Leave missing evidence and caption review pending. Preserve originals and stop if a required input or review is missing.
>
> Original prompt for an assistant with local tool access

## Frequently asked questions

Can I add voiceover to an existing screen recording?

Yes. Select a recording that visibly supports the narration, trim and crop a copy, and mute its original audio when adding the WAV. Preserve the source. This kit requires enough video to cover the rounded narration timeline.

Are the captions generated automatically by the kit?

No. The kit consumes explicit reviewed JSON cues. Local transcription can draft words and times, but someone must verify them against the take and finished movie. The supplied review flag is a record, not an independent listening check.

Does this work entirely offline?

Assembly uses local media and a cached browser. Dependencies and speech/transcription models may need initial downloads. Murmur also requires periodic online license validation. A hosted assistant can have its own network use even when speech generation runs locally.

What changes when I rewrite the narration?

Generate a new take, measure its duration and hash, rebuild the cue record, and repeat timing review. Prepare a new run and output. The kit rejects old caption metadata instead of assuming that a revised line has the same timing.

## Sources

- [Apple: screen recording in QuickTime Player](https://support.apple.com/en-au/guide/quicktime-player/qtp97b08e666/mac)Accessed 2026-10-09
- [Remotion: local Video component](https://www.remotion.dev/docs/media/video)Accessed 2026-10-09
- [Remotion: local Audio component](https://www.remotion.dev/docs/media/audio)Accessed 2026-10-09
- [Remotion: caption workflows](https://www.remotion.dev/docs/captions/)Accessed 2026-10-09
- [whisper.cpp: local transcription and CLI input format](https://github.com/ggml-org/whisper.cpp)Accessed 2026-10-09
- [W3C WAI: captions and automatic-caption accuracy](https://www.w3.org/WAI/media/av/captions/)Accessed 2026-10-09
- [W3C WAI: description of important visual information](https://www.w3.org/WAI/media/av/description/)Accessed 2026-10-09
- [Remotion: license and pricing guidance](https://www.remotion.dev/docs/license/pricing)Accessed 2026-10-09
- [Murmur: Mac requirements, website-edition offer, and license validation](https://www.murmurtts.com/)Accessed 2026-10-09

## Create your demo narration on your Mac

Generate local narration on an Apple Silicon Mac with macOS 15 or later. Murmur’s website edition is $49 one-time, with no free trial and a 7-day refund policy. Optional models need additional downloads and at least 16 GB of unified memory. Check the current requirements and listen to samples before buying.

[Get Murmur](https://murmur-licenses.tarunyadav9761.workers.dev/checkout)[Download Murmur](https://murmur-updates.tarunyadav9761.workers.dev/download/latest)
