Create a Promo Reel with AI and Local Mac Voiceover
Build a Claude Code video workflow: write a promo script, generate local Mac narration with Murmur, and render a vertical reel. Download tested commands.
Direct answer: Create a promo reel by asking a coding assistant to prepare the script and motion layout, generating the narration locally in Murmur, and rendering the pictures and audio into a vertical MP4. A Claude Code video workflow can keep the script, scene files, and render instructions in one project. Murmur supplies the voice track on your Mac; the assistant prepares the writing and code. This guide is for app builders and creators who want an editable reel rather than a single finished clip they cannot revise.
We tested the local part on October 9, 2026: one original script became a 7.12-second WAV through the Murmur CLI, then a 1080 by 1920 MP4 through FFmpeg. We also submitted a separate short take through MCP and retrieved its completed audio with the job tool. Download the tested commands and editable agent prompt to reproduce the small example before adapting it to your product. The installed helper came from a development build reporting Murmur 1.0.18; parity with the released binary was not verified.
Plan a Claude Code video around one idea
A short promo needs one visible result. Describe the viewer, the task your product solves, and the exact action at the end. For a first reel, a useful brief is a hook, a demonstrated result, and a next step. Decide which screenshots or original graphics will prove that result before asking for motion. If the product has no footage of a claimed feature, adjust the claim or capture the feature first.
The prompt below can be adapted for a coding assistant using Claude Opus 5.5, or an agent host configured with GPT-6 Astra. Those links confirm the model names. Neither model was run for this guide's fixture. The measured results describe local speech and rendering, not a comparison of assistant performance or proof that either model produces a complete video from one prompt.
Create a vertical promo reel for [product] aimed at [viewer]. Show [one verified result] and end with [one next step]. First write a short script and a three-scene plan. Keep narration separate from on-screen text. Use only the screenshots and original assets in this workspace. Check Murmur status, installed models, and preset voices before generating. Use one model and voice across scenes, save new WAV takes under audio/, measure the finished audio, and time the visuals from those measurements. Render a 1080x1920 MP4 at 30 fps. Return the script, editable source, audio paths, render command, and verified output metadata. Keep publication for my review.
Editable prompt for an agent with local file and tool access
Replace every bracketed field with a fact about your own product. A model name does not give the assistant access to a Mac, Murmur, screenshots, or a renderer. Your agent host must expose the relevant terminal or MCP tools. Keep the product demo script template beside the brief if the reel needs to explain an actual workflow rather than introduce a brand.
Check the Mac, model, and preset voice
Install Murmur on an Apple Silicon Mac running macOS 15 or later, finish setup, and enable local automation in Settings. Install the CLI from the same pane. The Murmur automation documentation covers the setup and agent connections. The video portion of this example separately needs FFmpeg, ffprobe, and Python 3 with Pillow; those are tools for the original title cards and movie render, not a manual installation of Murmur's speech runtime.
murmur --help
murmur status --json
murmur models --json
murmur voices --model qwen3-base --json
ffmpeg -version
ffprobe -version
python3 -c "import PIL; print(PIL.__version__)"Check that setup, licensing, and automation are ready, then choose a model listed as installed. On the test Mac, Qwen3-TTS Base was installed and exposed the English preset voices Ryan and Aiden. Kokoro was listed as unavailable there, so we used qwen3-base and Ryan explicitly. Your installed models can differ. Do not combine a voice ID copied from one model with another model merely because it is the current default.
Use the same preset for the first few edits. That makes it easier to decide whether the problem is the sentence, the pace, or the image timing. A saved reference voice is optional and is unnecessary for this example. Start with original narration and preset voices so the basic production loop can be checked before adding a more complicated performance requirement.
Generate an original local narration
Create a new project folder and save only spoken words in script.txt. Instructions such as camera movement, logo placement, or scene numbers belong in your visual plan. The original fixture uses this line: “Start with one clear idea. Write a short script. Generate the voice on your Mac, then build the motion around it.” It describes the production sequence and makes no claim about a fictional customer or product outcome.
murmur generate --input script.txt \
--model qwen3-base --voice Ryan \
--language EN-US --speed 1.0 \
--output narration.wav --jsonThis exact generation request completed with a succeeded state and an artifact identifying Qwen3-TTS Base, Ryan, and a duration of 7.12 seconds. ffprobe reported mono PCM audio at 24,000 Hz. That is a measurement of one short file on one Mac. A different script, model, voice, or generation can have a different duration, so use the finished file rather than copying the fixture's timing into a longer reel.
The CLI waits by default. Add --detach when you want a job ID immediately, then check that job before using its output. Keep --force absent while creating fresh takes; Murmur rejects an existing output unless replacement was requested. When changing copy, write narration-v2.wav and compare it with the previous take before updating the movie.
murmur generate --input script.txt --model qwen3-base \
--voice Ryan --output narration-v2.wav --detach --json
murmur jobs get JOB_ID --jsonBuild a simple vertical motion reel
The downloadable recipe makes three original title cards: write the script, generate the voice, and build the motion. Each card uses the same colors and typography, with a restrained zoom and a short fade. It reads the WAV duration, divides the picture sequence into three parts, and joins them before encoding. These are original graphics for a workflow example; they are not a recording of Murmur's interface or a finished advertisement for your own app.
Run the title-card and render helper in the fresh project folder after generating the WAV. The download contains the full Python helper and its FFmpeg arguments. It uses the same filters and encoding settings as the tested render, with paths relative to your folder. Pillow creates the PNGs, while FFmpeg performs the motion, concatenation, audio mapping, and MP4 encoding. Keep the editable text and images so a changed sentence can be repaired without starting again.
python3 make-reel.py
ffprobe -v error -count_frames -show_format -show_streams \
-of json reel-smoke-test.mp4
ffmpeg -hide_banner -loglevel error -i reel-smoke-test.mp4 \
-f null -The tested file contained H.264 video and AAC audio, measured 1080 by 1920 pixels at 30 frames per second, and decoded successfully from beginning to end. It contained 214 video frames and lasted 7.133333 seconds; the audio lasted 7.12 seconds. The small difference follows the video frame boundary. Its size was 210,042 bytes because the pictures were simple title cards. That size is not a prediction for screenshot-heavy or camera-based reels.
The official FFmpeg documentation explains input ordering and explicit stream selection. The filter reference covers the zoom and pan options used here. In the recipe, -map selects the generated picture stream and the narration, while -n prevents the renderer from replacing an existing movie. Inspect the output even when the process exits successfully.
Let MCP track the audio while you edit
MCP is useful when you want the agent to inspect models and submit speech work directly. Start the server from the video project folder so its file operations stay in that workspace. The documented Claude Code connection uses the helper shipped inside the installed app. For a typical installation, register it with the command below, then open the agent session in your project folder.
claude mcp add --transport stdio --scope user murmur -- \
/Applications/Murmur.app/Contents/Helpers/murmur mcp serveThe installed server exposed ten tools in our check. For this workflow, the useful sequence is murmur_status, murmur_list_models, murmur_list_voices, murmur_generate_speech, and murmur_get_job. These are the actual tool names, not illustrative API names. Speech generation returns a queued job by default, allowing the assistant to prepare pictures while the app produces audio.
{
"name": "murmur_generate_speech",
"arguments": {
"text": "Build a reel with local narration.",
"output_path": "mcp-take.wav",
"model_id": "qwen3-base",
"voice_id": "Ryan",
"language": "EN-US",
"speed": 1.0
}
}We submitted those arguments to the local stdio server and received a queued job. A later murmur_get_job call with its job_id returned succeeded and a 2.24-second WAV artifact. A queued response alone does not establish that usable audio exists. Wait for the terminal state and artifact path, then check the file. The download shows the job-polling arguments and a batch example for separate scene takes.
Choose what to revise after the first render
| What you hear or see | Change first | Keep for comparison |
|---|---|---|
| The line runs past the scene | Shorten the sentence or extend that scene | The original WAV and scene timing |
| The words are clear but feel rushed | Adjust pace, then measure the new take | The same model and preset voice |
| A product claim has no matching picture | Capture the actual result or rewrite the claim | The approved brief and source assets |
| Only one scene needs new copy | Generate a new scene take and update its timing | Unchanged scene audio |
| The movie exists but has missing audio | Check the mapped streams and decode the output | The render command and ffprobe report |
Listen to the exported movie on the device where you will review it, and watch every word of the title cards at normal playback speed. Our checks establish generated files, codec metadata, complete decoding, and a visually inspected midpoint frame. They do not establish that every spoken word was independently transcribed, that the performance is ideal, or that an audience will respond to the reel. Listening and editorial approval are still part of your production pass.
Add captions from the final voice track and verify them against the speech. The title cards in this fixture summarize the workflow; they are not timed word-for-word captions. If your product reel needs screenshots, keep controls and meaningful results readable long enough for a viewer to understand them. The speech time calculator helps with an early draft, while the voiceover export guide helps you choose master and delivery audio.
What unlimited local TTS means for reel production
For a licensed Murmur installation, local speech generation has no per-character fees or speech credit balance. You can revise a line and make another local take without buying additional TTS credits. The practical limits are your Mac's available memory, compute time, disk space, and the capabilities of the selected model. A local app does not make a very long script instantaneous or make every voice setting valid for every engine.
Your coding assistant is a separate part of the workflow. A hosted assistant can have subscription limits, API usage charges, and network requirements even while Murmur generates speech locally. Likewise, the sample renderer and its dependencies are separate from the Murmur app. Evaluate the whole setup for the work you actually produce, and use original or authorized material for your finished reel. Begin with one small example and check the outputs before expanding to a campaign of multiple videos.
Frequently asked questions
Can Claude Code create a promotional video with a local voiceover?
An agent with file and tool access can prepare a script, generate motion code, call Murmur through the CLI or MCP, and run a renderer. This guide verified the local speech and FFmpeg steps. It did not run Claude Opus 5.5 or measure how well Claude Code produces the full reel.
Which model should I use for AI voiceover for reels on Mac?
Choose an installed speech model and a preset voice from that model, then audition a short line. This fixture used Qwen3-TTS Base with Ryan because they were available on the test Mac. Keep the same pair while revising the script, and measure the duration of each finished take.
Does Murmur render the video itself?
Murmur produces the local voice audio. In this example, Pillow created original title cards and FFmpeg animated and encoded them with the narration. You can use another video editor or renderer, but the tool must still combine the approved visuals, audio, scene timing, and final export settings.
Is the entire AI reel workflow unlimited and free?
Murmur local speech does not use per-character credits on a licensed installation. The website edition is a paid app. Hosted coding assistants can have their own charges and usage limits, and your Mac has finite resources. The unlimited local TTS description applies to speech usage, not every service involved.
How do I know an MCP voiceover job has finished?
Read the job ID returned by murmur_generate_speech, then call murmur_get_job with job_id. Use the file only after the snapshot reports succeeded and lists its artifact path. Check the actual audio before rendering. A queued job is a submitted request and does not prove that speech generation completed.
Sources
- Murmur: CLI and MCP automation documentationAccessed 2026-10-09
- FFmpeg: inputs, stream selection, and output optionsAccessed 2026-10-09
- FFmpeg: filter reference, including zoompan and concatAccessed 2026-10-09
- Anthropic: Claude Opus 5.5 announcement, model namingAccessed 2026-10-09
- OpenAI: GPT-6 Astra model page, model namingAccessed 2026-10-09