PROMO REEL WITH LOCAL MAC VOICEOVER Checked October 9, 2026 Article: /blog/resources/ai-promo-reel-local-voiceover-mac This original short fixture verifies Murmur CLI speech, MCP speech/job retrieval, and a local FFmpeg vertical render. Claude Opus 5.5 and GPT-6 Astra are suggested assistant choices for adapting the prompt; neither was run for this fixture. Murmur generates voice audio. The assistant writes/plans code and assets. REQUIREMENTS Apple Silicon Mac, macOS 15+, licensed Murmur with setup complete and local automation enabled. Install CLI in Murmur Settings > Automation. The video helper separately needs ffmpeg, ffprobe, Python 3 and Pillow. Test environment: development build reporting Murmur 1.0.18; FFmpeg 9.0.2; Pillow 11.3.0; Apple M3 Pro. Released-binary parity was not verified. If these video dependencies are absent, configure them separately or use your existing video editor. They are not prerequisites for Murmur's managed TTS runtime. Use a fresh project folder, and pick an installed model/voice from your own Mac. 1. CHECK READINESS murmur --help murmur status --json murmur models --json murmur voices --model qwen3-base --json ffmpeg -version ffprobe -version python3 -c "import PIL; print(PIL.__version__)" The fixture used qwen3-base / Ryan. The Mac also exposed the English Aiden preset. Kokoro was unavailable on this Mac, so do not assume af_bella works with the current default model. Never change defaults just to run this fixture. 2. CREATE ORIGINAL SCRIPT IN A NEW PROJECT FOLDER mkdir promo-reel-01 cd promo-reel-01 cat > script.txt <<'SCRIPT' Start with one clear idea. Write a short script. Generate the voice on your Mac, then build the motion around it. SCRIPT 3. GENERATE A LOCAL WAV murmur generate --input script.txt --model qwen3-base --voice Ryan \ --language EN-US --speed 1.0 --output narration.wav --json The original result was succeeded, with Qwen3-TTS Base / Ryan, 7.12 seconds, mono 24,000 Hz PCM signed 16-bit WAV, 341,804 bytes. Measure your own result. No --force is used. For a revision choose a new output filename. Optional detached request (separate new take; not required for the render): murmur generate --input script.txt --model qwen3-base --voice Ryan \ --output narration-v2.wav --detach --json murmur jobs get JOB_ID --json Replace JOB_ID with the returned UUID. Wait for succeeded plus an artifact path. 4. SAVE THE ORIGINAL TITLE-CARD / FFMPEG HELPER cat > make-reel.py <<'PY' from PIL import Image, ImageDraw, ImageFont from pathlib import Path import subprocess, json, math root = Path.cwd() targets = [root / f'slide-{n}.png' for n in range(1, 4)] targets += [root / name for name in ('reel-smoke-test.mp4', 'render-command.json', 'output-metadata.json')] if any(path.exists() for path in targets): raise SystemExit('Output exists. Use a fresh folder or new filenames for this version.') font_path = '/System/Library/Fonts/Supplemental/Arial.ttf' font = ImageFont.truetype(font_path, 86) small = ImageFont.truetype(font_path, 42) slides = [ ('WRITE THE SCRIPT', 'Start with one clear idea.'), ('GENERATE THE VOICE', 'Local narration on your Mac.'), ('BUILD THE MOTION', 'Let the audio set the pace.') ] for n, (title, subtitle) in enumerate(slides, 1): im = Image.new('RGB', (1080, 1920), (17, 24, 39)) d = ImageDraw.Draw(im) d.rounded_rectangle((84, 500, 996, 1180), radius=44, fill=(24, 33, 54)) d.rectangle((84, 500, 996, 516), fill=(145, 162, 255)) d.text((118, 584), f'0{n}', font=font, fill=(145, 162, 255)) for j, word in enumerate(title.split()): d.text((118, 742 + j*110), word, font=font, fill=(245, 247, 255)) d.text((118, 1110), subtitle, font=small, fill=(191, 199, 224)) d.text((118, 1350), 'Murmur workflow smoke test', font=small, fill=(145, 162, 255)) im.save(root / f'slide-{n}.png') audio = root / 'narration.wav' metadata = json.loads(subprocess.check_output([ 'ffprobe', '-v', 'error', '-show_format', '-show_streams', '-of', 'json', str(audio) ], text=True)) duration = float(metadata['format']['duration']) parts = math.ceil(duration * 30 / 3) / 30 inputs = [] for n in range(1, 4): inputs += ['-loop', '1', '-framerate', '30', '-i', str(root / f'slide-{n}.png')] filters = [] for n in range(3): filters.append( f"[{n}:v]zoompan=z='min(zoom+0.00035,1.03)':d=1:s=1080x1920:fps=30," f"trim=duration={parts},fade=t=in:st=0:d=0.15," f"fade=t=out:st={parts-0.15}:d=0.15,setpts=PTS-STARTPTS[v{n}]" ) filters.append('[v0][v1][v2]concat=n=3:v=1:a=0,format=yuv420p[v]') command = [ 'ffmpeg', '-hide_banner', '-loglevel', 'warning', '-n', *inputs, '-i', str(audio), '-filter_complex', ';'.join(filters), '-map', '[v]', '-map', '3:a:0', '-t', str(duration), '-c:v', 'libx264', '-preset', 'fast', '-crf', '23', '-c:a', 'aac', '-b:a', '192k', '-ar', '48000', '-movflags', '+faststart', str(root / 'reel-smoke-test.mp4') ] (root / 'render-command.json').write_text(json.dumps(command, indent=2)) subprocess.run(command, check=True) output = json.loads(subprocess.check_output([ 'ffprobe', '-v', 'error', '-count_frames', '-show_format', '-show_streams', '-of', 'json', str(root / 'reel-smoke-test.mp4') ], text=True)) (root / 'output-metadata.json').write_text(json.dumps(output, indent=2)) print(json.dumps(output, indent=2)) PY python3 make-reel.py This uses the same FFmpeg arguments/filters as the measured fixture. The helper uses current-folder paths instead of the test's temporary folder. The helper checks every generated filename before creating cards, and FFmpeg -n refuses to replace a movie. For another version use a new folder or change all output filenames. Existing audio is used as input and is never replaced. 5. CHECK THE FINAL MP4 ffprobe -v error -count_frames -show_format -show_streams \ -of json reel-smoke-test.mp4 ffmpeg -hide_banner -loglevel error -i reel-smoke-test.mp4 -f null - Original measured video: 1080x1920, 30 fps, H.264 High/yuv420p, AAC mono 48,000 Hz, 214 video frames, 7.133333 seconds, 210,042 bytes. Full video/audio decode returned 0. A midpoint frame was visually inspected. The title cards are summaries, not timed speech captions. Listen to the completed movie, confirm the words, review timing, and add checked captions before using it as a public reel. This test does not independently transcribe every spoken word or measure audience response. Your duration and size may differ; use the actual metadata. 6. MCP OPTION: LET AN AGENT REQUEST AUDIO For a normal app installation, the documented Claude Code registration is: claude mcp add --transport stdio --scope user murmur -- \ /Applications/Murmur.app/Contents/Helpers/murmur mcp serve For Codex: codex mcp add murmur -- /Applications/Murmur.app/Contents/Helpers/murmur mcp serve Open the agent in this project folder. Call murmur_status, murmur_list_models, and murmur_list_voices before generating. The MCP server limits file access to the directory where it launches. Typical tool-call parameters: { "name": "murmur_generate_speech", "arguments": { "text": "Build a reel with local narration.", "output_path": "mcp-take.wav", "model_id": "qwen3-base", "voice_id": "Ryan", "language": "EN-US", "speed": 1.0 } } This exact local tool invocation was tested and returned a queued job. Poll with: { "name": "murmur_get_job", "arguments": { "job_id": "RETURNED_JOB_UUID" } } The original completed artifact was Qwen3-TTS Base / Ryan, 2.24 seconds. MCP generation is asynchronous by default; "wait": true requests completion before returning. A queued response alone is not audio completion. Batch shape verified in the installed tool schema, not executed for this fixture: { "name": "murmur_generate_batch", "arguments": { "items": [ { "id": "hook", "text": "Start with one clear idea.", "output_path": "audio/01-hook.wav", "model_id": "qwen3-base", "voice_id": "Ryan", "speed": 1.0 }, { "id": "result", "text": "Build the motion around your voiceover.", "output_path": "audio/02-result.wav", "model_id": "qwen3-base", "voice_id": "Ryan", "speed": 1.0 } ] } } CLI batch manifests instead use camelCase field names, such as outputPath, modelID and voiceID, then: murmur batch --manifest scenes.json --json 7. EDITABLE AGENT PROMPT Create a vertical promo reel for [product] aimed at [viewer]. Show [one verified result] and end with [one next step]. First write a short script and a three-scene plan. Keep narration separate from on-screen text. Use only the screenshots and original assets in this workspace. Check Murmur status, installed models, and preset voices before generating. Use one model and voice across scenes, save new WAV takes under audio/, measure the finished audio, and time the visuals from those measurements. Render a 1080x1920 MP4 at 30 fps. Return the script, editable source, audio paths, render command, and verified output metadata. Keep publication for my review. Adapt that prompt to your agent host and selected coding model. We checked the provider pages for Claude Opus 5.5 and GPT-6 Astra naming, but did not run either for this sample. The host needs local file/terminal access or Murmur MCP tools. Local TTS has no per-character fees on a licensed Murmur installation. Your assistant plan/API charges and finite Mac resources remain separate. CHECKED PRIMARY REFERENCES https://www.murmurtts.com/docs/automation https://ffmpeg.org/ffmpeg.html https://ffmpeg.org/ffmpeg-filters.html https://www.anthropic.com/claude-opus-5-5 https://developers.openai.com/api/docs/models/gpt-6-astra